Senior Software Engineer - Model Evaluation & AI Systems
Core
Build automated evaluation pipelines and infrastructure to validate the quality of speech-to-text, text-to-speech, and multimodal AI models before they reach customers.
Role type
Senior IC software engineer (AI evaluation & systems)
Builds
Automated evaluation pipelines, canaries, and continuous-monitoring systems for production AI models
Domain
Voice AI, Speech-to-Text, Text-to-Speech, LLMs, Multimodal systems
Deliverable
production ML models
Required skills
Python, Rust, Go, automated test pipelines, data processing systems, statistical analysis, CI/CD integration
Preferred skills
LLM/RAG/agent evaluation, React Native, ML infrastructure, voice/audio metrics (WER, MOS), cloud infrastructure, containerized environments
Technologies
Python, Rust, Go, Grafana, CI/CD tools
Responsibilities
Define evaluation methodologies for STT, TTS, and emerging AI systems; Design and maintain automated evaluation pipelines for batch and streaming; Build scalable evaluation infrastructure running on production models and GPU clusters; Translate research benchmarks into automated pass/fail gates; Operate canaries and monitoring systems to detect quality regressions; Partner with DevOps to stand up ephemeral test environments; Integrate quality gates into CI/CD workflows
Seniority
Senior, hands-on IC