CareerPlanSign in

ML Engineer, Inference & Optimization

Palo Alto HQ💼 Full-time🗓 2026-06-23 → 2026-09-25

Core

Accelerate inference performance and efficiency for Pika's AI-driven video and language models through GPU parallelism and advanced deployment techniques.

Role type

Senior/Staff Inference Engineer (Video & LLMs)

Builds

High-performance inference pipelines for real-time video generation and large language models

Domain

AI Infrastructure / Video Generation / Large Language Models

Deliverable

production ML models

Required skills

Inference optimization, Quantization, Attention acceleration, Deep learning compiler stacks, GPU programming (CUDA, NCCL), Distributed parallelism (TP, SP, PP), High-performance computing kernels

Preferred skills

Training efficiency optimization, Distributed training toolkits, Open source contributions to AI infrastructure, Startup prototyping experience

Technologies

CUDA, NCCL, Tensor Parallelism, Sequence Parallelism, Pipeline Parallelism

Responsibilities

Implement attention optimization and quantization for efficient model serving, Engineer GPU strategies for maximal efficiency and scalability, Develop high-performance computing kernels and distributed workloads, Collaborate with research teams to deploy state-of-the-art models into production, Mentor engineers on inference and GPU programming best practices

Seniority

Senior/Staff, hands-on IC with mentorship

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.