Speech Evaluation Lead
Core
Designing and scaling evaluation pipelines, human studies, and automated metrics to measure the quality, naturalness, and safety of large-scale speech and multimodal AI models.
Role type
Founding evaluation lead for speech and multimodal models
Builds
Evaluation pipelines, automated dashboards, and auxiliary models for TTS, voice conversion, and ASR systems
Domain
AI/ML, Speech Technology, Multimodal Systems
Deliverable
production ML models
Required skills
Building evaluation pipelines for speech/multimodal models, designing human studies (MOS, ABX), defining objective metrics (WER, intelligibility, prosody), Python, ML frameworks (PyTorch, Hugging Face), statistics and experimental design
Preferred skills
Experience with ASR, TTS, or voice cloning pipelines, cross-functional collaboration
Technologies
Python, PyTorch, Hugging Face
Responsibilities
Build and scale evaluation pipelines for TTS, voice conversion, and ASR systems; Design human studies for subjective testing; Define and implement objective metrics; Automate evaluation dashboards and reporting systems; Train auxiliary models to capture new evaluation dimensions; Collaborate across data, model, and product teams to drive measurable improvement; Establish and scale the evaluation function as the team grows
Seniority
Founding, hands-on IC with leadership scope