Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco
Core
Building high-throughput, ultra-low-latency inference engines for large language models and foundational speech models to power real-time conversational AI.
Role type
Senior IC machine learning engineer (inference & serving)
Builds
Real-time speech LLM inference systems for hardware-software AI companions
Domain
AI/ML, Speech Technology, Distributed Systems
Deliverable
production ML models
Required skills
GPU architecture optimization, continuous batching, KV cache management, real-time audio streaming, model compression & quantization, distributed inference pipelines
Preferred skills
vLLM, TensorRT-LLM, SGLang, NVIDIA Triton Inference Server, speculative decoding, chunked prefill, Kubernetes autoscaling
Technologies
NVIDIA Ampere/Hopper GPUs, WebSockets, WebRTC, Kubernetes, FP8, INT8, AWQ, GPTQ
Responsibilities
Deploy multi-GPU and multi-node inference pipelines, manage autoscaling infrastructure, optimize hardware bottlenecks, implement advanced generation algorithms, handle continuous audio streams
Seniority
Senior, hands-on IC