CareerPlanSign in

Data Scientist, Agent Evaluations & Quality

Palo Alto💼 Full-time🗓 2026-09-24 → 2026-09-25

Core

Architecting automated evaluation pipelines and metrics frameworks to measure the quality of autonomous AI agents handling email, calendar, and business software tasks.

Role type

Senior IC data scientist (AI agent evaluation & quality)

Builds

Automated evaluation pipelines, gold datasets, regression suites, deterministic/model-based graders, and release-quality dashboards

Domain

AI productivity, autonomous agents, LLM systems

Deliverable

production ML models | dashboards & analysis

Required skills

evaluation system design, metrics framework development, Python, SQL, statistical experimental design, ground-truth data development, LLM behavior analysis, failure taxonomy construction, stakeholder communication

Preferred skills

LLM-as-a-judge systems, agentic task evaluation, benchmarking platforms

Technologies

Python, SQL, LLMs, LLM-as-a-judge

Responsibilities

Architect automated evaluation pipelines; define success criteria for multi-step tasks; build gold datasets and regression suites; design graders and calibrate LLM-as-a-judge systems; analyze traces to identify root causes; compare models and prompts via offline experiments; build actionable dashboards for engineering and product teams

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.