Senior Research Engineer - Voice
Core
Developing high-quality, expressive, and real-time synthetic voices for an AI video platform used by Fortune 100 companies.
Role type
Senior Research Engineer (Generative Audio/Speech)
Builds
Generative speech and voice synthesis models for enterprise skill development and visual communication.
Domain
Generative AI, Audio/Speech Technology, Large Language Models
Deliverable
production ML models
Required skills
Generative modeling, Large Language Models (LLMs), PyTorch, Time-series modeling, Tokenization, Deep learning model training, Software engineering
Preferred skills
Streaming architectures, Audio diffusion models, Neural codecs, Flow-matching models, Speech-to-speech systems, Academic publications
Technologies
PyTorch, LLMs, Diffusion models, Neural codecs, Flow-matching models
Responsibilities
Develop and evaluate streaming and speech-to-speech systems; Adapt models for new conditioning inputs (emotion, speed, prosody); Implement post-training optimization techniques (quantization, pruning, distillation); Integrate and test novel architectures (neural codecs, diffusion, flow-matching); Define new evaluation metrics for conversational speech; Apply DPO and distillation to fine-tune large-scale speech models.
Seniority
Senior, hands-on IC