CareerPlanGet AI match score →

Senior Machine Learning Engineer, Voice AI

San Francisco💼 Full-time💰 $200,000–$200,000🗓 2026-07-10 → 2026-07-31

Core

Building the model serving layer for production-grade, real-time voice agents and applications (STT, TTS, speech-to-speech) with best-in-class latency and reliability.

Role type

Senior IC machine learning engineer (voice inference)

Builds

Inference infrastructure for voice models (Whisper, Parakeet, Orpheus, Kokoro) and partners (Cartesia, Deepgram, Rime)

Domain

AI infrastructure / Voice AI / Real-time streaming

Deliverable

production ML models

Required skills

LLM serving engines (vLLM, SGLang, TensorRT-LLM), Python, PyTorch, GPU profiling and optimization (CUDA, memory management, kernel-level debugging), streaming audio handling, model evaluation frameworks

Preferred skills

Speech and audio ML (ASR, TTS architectures, audio signal processing), audio codecs and tokenization schemes (SNAC, Encodec, DAC), model fine-tuning

Technologies

TRT-LLM, SGLang, H100s, H200s, B200s, vLLM, TensorRT-LLM, CUDA, PyTorch

Responsibilities

Optimize inference performance for voice models targeting best-in-class TTFB, throughput, and GPU utilization; Productionize voice models on serverless and dedicated endpoints with batching strategies and streaming inference; Build and maintain a voice model evaluation framework measuring WER, naturalness, and latency; Enable new model architectures including audio-native LLMs and codec-based models; Profile and debug performance across the full inference stack from GPU kernels to framework-level bottlenecks; Collaborate with platform engineering to meet latency and reliability requirements for real-time voice APIs

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗