Staff / Principal Machine Learning Engineer, Serving - USA
Core
Building and optimizing high-performance, sub-second multimodal inference systems for real-time voice models serving hundreds of millions of users across consumer AI applications.
Role type
Staff / Principal Machine Learning Engineer (Inference Optimization & Serving)
Builds
Realtime voice models, optimized inference engines, and scalable serving APIs
Domain
AI / Machine Learning / Real-time Inference
Deliverable
production ML models
Required skills
Inference optimization (vLLM, TRT-LLM), Model acceleration (quantization, distillation, caching, continuous batching, paged attention, speculative decoding), High-performance systems programming (C++, CUDA, Rust, optimized Python), Distributed systems & scaling (Kubernetes, Ray, multi-GPU/multi-node inference), Full-cycle model ownership
Preferred skills
Open-source contributions to major inference engines, Deep-dive technical write-ups, PhD in CS/Physics/Math or equivalent practical experience
Technologies
vLLM, TRT-LLM, Kubernetes, Ray, C++, CUDA, Rust, Python, NVIDIA GPUs
Responsibilities
Optimize model inference latency and throughput, implement advanced serving techniques, profile and tune code for maximum GPU performance, manage distributed inference clusters, ensure production reliability of voice models
Seniority
Staff / Principal, hands-on IC with strategic impact