Lead Machine Learning Engineer, Evaluations
Core
Design and own the evaluation platform to measure quality, safety, and performance of ASAPP's agentic AI systems and LLM-based solutions.
Role type
Lead Machine Learning Engineer (Evaluation Infrastructure)
Builds
Evaluation platform infrastructure, benchmarking pipelines, and monitoring systems for agentic AI
Domain
AI Engineering / NLP / Agentic Systems / Customer Experience
Deliverable
production ML models | infrastructure
Required skills
evaluation system design, architectural leadership, Python, AWS, Kubernetes, Docker, data pipeline design, dataset versioning, LLM-as-judge methodologies, human-in-the-loop workflows
Preferred skills
agentic systems at scale, voice/audio quality evaluation, LLM-centric services, large-scale ML experimentation, conversational AI domains, model optimization for inference, CI/CD, Kafka, Athena
Responsibilities
Develop technical roadmap and architecture for the evaluation platform; Design eval methodologies including golden/regression test sets and automated metrics; Build data infrastructure for annotation, labeling, and dataset versioning; Partner with Research, Product, and Platform teams to productize experiments; Represent the eval platform to stakeholders and report on platform health; Mentor and support other engineers through design reviews and knowledge sharing
Seniority
Lead, hands-on IC with mentorship responsibilities