Data Scientist, Agent Evaluations & Quality
Core
Architecting automated evaluation pipelines and metrics frameworks to measure the quality of autonomous AI agents handling email, calendar, and business software tasks.
Role type
Senior IC data scientist (AI agent evaluation & quality)
Builds
Automated evaluation pipelines, gold datasets, regression suites, deterministic/model-based graders, and release-quality dashboards
Domain
AI productivity, autonomous agents, LLM systems
Deliverable
production ML models | dashboards & analysis
Required skills
evaluation system design, metrics framework development, Python, SQL, statistical experimental design, ground-truth data development, LLM behavior analysis, failure taxonomy construction, stakeholder communication
Preferred skills
LLM-as-a-judge systems, agentic task evaluation, benchmarking platforms
Technologies
Python, SQL, LLMs, LLM-as-a-judge
Responsibilities
Architect automated evaluation pipelines; define success criteria for multi-step tasks; build gold datasets and regression suites; design graders and calibrate LLM-as-a-judge systems; analyze traces to identify root causes; compare models and prompts via offline experiments; build actionable dashboards for engineering and product teams
Seniority
Senior, hands-on IC
