Senior Machine Learning Engineer, Voice AI
Core
Building the model serving layer for production-grade, real-time voice agents and applications (STT, TTS, speech-to-speech) with best-in-class latency and reliability.
Role type
Senior IC machine learning engineer (voice inference)
Builds
Inference infrastructure for voice models (Whisper, Parakeet, Orpheus, Kokoro) and partners (Cartesia, Deepgram, Rime)
Domain
AI infrastructure / Voice AI / Real-time streaming
Deliverable
production ML models
Required skills
LLM serving engines (vLLM, SGLang, TensorRT-LLM), Python, PyTorch, GPU profiling and optimization (CUDA, memory management, kernel-level debugging), streaming audio handling, model evaluation frameworks
Preferred skills
Speech and audio ML (ASR, TTS architectures, audio signal processing), audio codecs and tokenization schemes (SNAC, Encodec, DAC), model fine-tuning
Technologies
TRT-LLM, SGLang, H100s, H200s, B200s, vLLM, TensorRT-LLM, CUDA, PyTorch
Responsibilities
Optimize inference performance for voice models targeting best-in-class TTFB, throughput, and GPU utilization; Productionize voice models on serverless and dedicated endpoints with batching strategies and streaming inference; Build and maintain a voice model evaluation framework measuring WER, naturalness, and latency; Enable new model architectures including audio-native LLMs and codec-based models; Profile and debug performance across the full inference stack from GPU kernels to framework-level bottlenecks; Collaborate with platform engineering to meet latency and reliability requirements for real-time voice APIs
Seniority
Senior, hands-on IC