Machine Learning Engineer, TTS
Core
Build state-of-the-art speech systems end-to-end, driving the model ↔ data ↔ eval flywheel for TTS and adjacent tasks like voice cloning and voice conversion.
Role type
Senior Research / ML Engineer (Speech/Audio)
Builds
Large-scale speech models (TTS, voice cloning, voice conversion) and production inference pipelines
Domain
AI, Speech Technology, Audio Modeling
Deliverable
production ML models
Required skills
Large-scale audio model development (>3B models, >500k hours data), Transformer architectures, Diffusion models, Multi-node/multi-GPU distributed training, PyTorch, CUDA/Triton/C++, Production code quality, Data curation and augmentation, Automated evaluation design, GPU scaling, Safety guardrails
Preferred skills
Voice-cloning, Speech-control, Audio language modelling, Open source contributions in speech/audio/ML
Technologies
PyTorch, CUDA, Triton, C++, GRPO, DPO
Responsibilities
Architect and implement large-scale speech models, Design and run scientific experiments, Develop dev tooling, Harden training/evaluation/inference pipelines, Profile latency/memory/cost, Contribute to safety and misuse mitigation
Seniority
Senior, hands-on IC with research leadership