CareerPlanSign in

Staff Machine Learning Engineer, Voice AI

San Francisco💼 Full-time🗓 2026-07-10 → 2026-09-26

Core

Building and optimizing the model serving layer for real-time voice AI applications (STT, TTS, speech-to-speech) to ensure best-in-class latency and throughput.

Role type

Staff Machine Learning Engineer (Voice Inference Infrastructure)

Builds

Production-grade inference engines and serving platforms for voice models (Whisper, Parakeet, Orpheus, Kokoro) on GPU clusters.

Domain

Voice AI / Real-time Inference / GPU Infrastructure

Deliverable

production ML models

Required skills

LLM serving engine internals (vLLM, SGLang, TensorRT-LLM), GPU optimization (CUDA, memory hierarchies), Python, PyTorch, system architecture, performance profiling, streaming audio pipelines

Preferred skills

Experience with audio-native LLMs, codec-based architectures (SNAC, Encodec), fine-tuning infrastructure

Technologies

TRT-LLM, SGLang, vLLM, H100/H200/B200 GPUs, PyTorch, CUDA

Responsibilities

Own the voice inference roadmap and technical strategy; architect high-performance serving systems for streaming audio; lead productionization of voice models at scale; build evaluation frameworks for model selection; enable next-generation model paradigms; drive technical leadership for model partner integrations; resolve complex performance bottlenecks; define scalable fine-tuning capabilities.

Seniority

Staff, high-impact technical leadership

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.