ML Engineer, Inference & Optimization
Core
Accelerate inference performance and efficiency for Pika's AI-driven video and language models through GPU parallelism and advanced deployment techniques.
Role type
Senior/Staff Inference Engineer (Video & LLMs)
Builds
High-performance inference pipelines for real-time video generation and large language models
Domain
AI Infrastructure / Video Generation / Large Language Models
Deliverable
production ML models
Required skills
Inference optimization, Quantization, Attention acceleration, Deep learning compiler stacks, GPU programming (CUDA, NCCL), Distributed parallelism (TP, SP, PP), High-performance computing kernels
Preferred skills
Training efficiency optimization, Distributed training toolkits, Open source contributions to AI infrastructure, Startup prototyping experience
Technologies
CUDA, NCCL, Tensor Parallelism, Sequence Parallelism, Pipeline Parallelism
Responsibilities
Implement attention optimization and quantization for efficient model serving, Engineer GPU strategies for maximal efficiency and scalability, Develop high-performance computing kernels and distributed workloads, Collaborate with research teams to deploy state-of-the-art models into production, Mentor engineers on inference and GPU programming best practices
Seniority
Senior/Staff, hands-on IC with mentorship