CareerPlanSign in

Senior Inference Engineer, AGI

Sunnyvale, California, United States💼 Full-time🗓 2026-08-27 → 2026-09-25

Core

Senior Inference Engineer owning the full-stack inference path for real-time multimodal conversational AI, co-designing model architectures for servability and building low-latency runtimes.

Role type

Senior Inference Engineer (Real-time Multimodal AI)

Builds

Real-time streaming serving runtime, offline training/RL infrastructure, and high-performance inference kernels.

Domain

AI/ML, Speech & Audio, Real-time Systems

Deliverable

production ML models

Required skills

Neural deep learning methods, Transformer architectures, GPU performance optimization, Inference optimization (quantization, speculative decoding, operator fusion), High-performance kernel development, Real-time streaming system design, Multi-GPU inference, Distributed training pipelines

Preferred skills

Custom GPU kernel authoring (CUTLASS, Triton, CUDA), Model compression techniques, Reinforcement learning infrastructure, Distributed training (NCCL, NVLink), Speech-to-speech/audio generative models, Open-source inference contributions

Technologies

vLLM, PyTorch, TensorRT-LLM, CUDA, CUTLASS, Triton, Nsight Compute, FlashAttention, NCCL, NVLink

Responsibilities

Partner with scientists to co-design inference-friendly model architectures; Implement and optimize inference paths for large-scale multimodal models; Develop high-performance kernels for critical operations; Own the real-time serving path for streaming multimodal conversational AI; Build offline inference systems for post-training and reinforcement learning; Profile end-to-end performance to eliminate bottlenecks.

Seniority

Senior, hands-on IC

Sourced via amazon · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.