CareerPlanSign in

Lead Machine Learning Engineer, Evaluations

Mountain View💼 Full-time🗓 2026-09-11 → 2026-09-26

Core

Design and own the evaluation platform to measure quality, safety, and performance of ASAPP's agentic AI systems and LLM-based solutions.

Role type

Lead Machine Learning Engineer (Evaluation Infrastructure)

Builds

Evaluation platform infrastructure, benchmarking pipelines, and monitoring systems for agentic AI

Domain

AI Engineering / NLP / Agentic Systems / Customer Experience

Deliverable

production ML models | infrastructure

Required skills

evaluation system design, architectural leadership, Python, AWS, Kubernetes, Docker, data pipeline design, dataset versioning, LLM-as-judge methodologies, human-in-the-loop workflows

Preferred skills

agentic systems at scale, voice/audio quality evaluation, LLM-centric services, large-scale ML experimentation, conversational AI domains, model optimization for inference, CI/CD, Kafka, Athena

Responsibilities

Develop technical roadmap and architecture for the evaluation platform; Design eval methodologies including golden/regression test sets and automated metrics; Build data infrastructure for annotation, labeling, and dataset versioning; Partner with Research, Product, and Platform teams to productize experiments; Represent the eval platform to stakeholders and report on platform health; Mentor and support other engineers through design reviews and knowledge sharing

Seniority

Lead, hands-on IC with mentorship responsibilities

Sourced via lever · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.