Sr. Machine Learning Engineer, Speech LLM Evaluation
Core
Owns the data and metrics foundation for evaluating speech LLMs (real-time speech understanding and generation models) across accuracy, robustness, and conversational quality.
Role type
Senior IC machine-learning engineer (speech LLM evaluation)
Builds
Evaluation datasets, metrics, and automated judges for audio LLMs
Domain
Consumer AI / Speech & Audio / Large Language Models
Deliverable
production ML models
Required skills
Python, data processing pipelines at scale, dataset curation/annotation, statistics for model performance, LLM evaluation techniques (automated and human), audio/speech evaluation
Preferred skills
audio-native or multimodal LLM evaluation, human evaluation study design, personalization and named-entity evaluation, multilingual audio dataset development, distributed data processing (Spark), publication record in speech/audio ML/NLP evaluation