CareerPlanSign in

Machine Learning Scientist

United States🌐 Remote💼 Full-time🗓 2026-05-15 → 2026-09-26

Core

Design, train, and evaluate speech synthesis and understanding models to build voice AI for enterprise customer experiences.

Role type

Senior IC machine learning scientist (speech synthesis/audio)

Builds

Production-grade text-to-speech and speech understanding models for high-volume conversational deployments

Domain

Voice AI, speech synthesis, audio processing

Deliverable

production ML models

Required skills

speech synthesis literature knowledge, neural codecs, multi-modal modeling, data quality assessment, TTS frontend integration, PyTorch, distributed training

Preferred skills

multilingual TTS, prosody/paralinguistics, published research, model quantization/distillation

Technologies

PyTorch, EnCodec, DAC, Mimi, Tacotron, FastSpeech, VITS, VALL-E, Moshi, LLaMA-Omni

Responsibilities

Design and train autoregressive and non-autoregressive speech synthesis models; Drive research on full-duplex and half-duplex multi-modal architectures; Iterate on speech representations including neural codecs and semantic tokens; Build rigorous objective and perceptual evaluation pipelines; Collaborate with linguists on TTS frontend behavior

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.