Machine Learning Scientist
Core
Design, train, and evaluate speech synthesis and understanding models to build voice AI for enterprise customer experiences.
Role type
Senior IC machine learning scientist (speech synthesis/audio)
Builds
Production-grade text-to-speech and speech understanding models for high-volume conversational deployments
Domain
Voice AI, speech synthesis, audio processing
Deliverable
production ML models
Required skills
speech synthesis literature knowledge, neural codecs, multi-modal modeling, data quality assessment, TTS frontend integration, PyTorch, distributed training
Preferred skills
multilingual TTS, prosody/paralinguistics, published research, model quantization/distillation
Technologies
PyTorch, EnCodec, DAC, Mimi, Tacotron, FastSpeech, VITS, VALL-E, Moshi, LLaMA-Omni
Responsibilities
Design and train autoregressive and non-autoregressive speech synthesis models; Drive research on full-duplex and half-duplex multi-modal architectures; Iterate on speech representations including neural codecs and semantic tokens; Build rigorous objective and perceptual evaluation pipelines; Collaborate with linguists on TTS frontend behavior
Seniority
Senior, hands-on IC