Machine Learning Engineer - Voice Conversion
Core
Build state-of-the-art speech systems end-to-end, driving the model ↔ data ↔ eval flywheel for Voice Conversion (VC) and adjacent tasks like controllable TTS and voice design.
Role type
Senior Research / ML Engineer (Speech/Audio)
Builds
Safe, steerable, and trustworthy AI speech systems for a social platform featuring lifelike AI characters.
Domain
Generative AI, Speech Technology, Audio Processing
Deliverable
production ML models
Required skills
Large-scale audio model development (>8B params, >500k hours), diffusion/flow-matching transformers, audio VAEs/neural codecs/vocoders, multi-node distributed training (FSDP/DeepSpeed), PyTorch, CUDA/Triton/C++, voice cloning, speech control/steerability
Preferred skills
Notable publications, open-source contributions in speech/audio/ML
Technologies
PyTorch, FSDP, DeepSpeed, CUDA, Triton, C++, GRPO, DPO
Responsibilities
Architect and train large-scale speech models; design and analyze scientific experiments; develop dev tooling; define data requirements and strategies; design automated objective/subjective evaluations; harden training/evaluation/inference pipelines; contribute to safety and misuse mitigation.
Seniority
Senior, hands-on IC