Staff Machine Learning Engineer, Voice AI
Core
Building and optimizing the model serving layer for real-time voice AI applications (STT, TTS, speech-to-speech) to ensure best-in-class latency and throughput.
Role type
Staff Machine Learning Engineer (Voice Inference Infrastructure)
Builds
Production-grade inference engines and serving platforms for voice models (Whisper, Parakeet, Orpheus, Kokoro) on GPU clusters.
Domain
Voice AI / Real-time Inference / GPU Infrastructure
Deliverable
production ML models
Required skills
LLM serving engine internals (vLLM, SGLang, TensorRT-LLM), GPU optimization (CUDA, memory hierarchies), Python, PyTorch, system architecture, performance profiling, streaming audio pipelines
Preferred skills
Experience with audio-native LLMs, codec-based architectures (SNAC, Encodec), fine-tuning infrastructure
Technologies
TRT-LLM, SGLang, vLLM, H100/H200/B200 GPUs, PyTorch, CUDA
Responsibilities
Own the voice inference roadmap and technical strategy; architect high-performance serving systems for streaming audio; lead productionization of voice models at scale; build evaluation frameworks for model selection; enable next-generation model paradigms; drive technical leadership for model partner integrations; resolve complex performance bottlenecks; define scalable fine-tuning capabilities.
Seniority
Staff, high-impact technical leadership