Quality Lead, Agentic AI Workflow Evaluation
Core
Senior individual contributor accountable for building and owning the quality system for evaluating complex, real-world agentic AI workflows in isolated test environments.
Role type
Senior individual contributor quality lead (agentic AI evaluation)
Builds
Quality audit systems, scoring standards, and evaluation frameworks for frontier AI agents
Domain
Artificial Intelligence / Agentic Systems / Quality Assurance
Deliverable
production ML models
Required skills
Quality assurance management, agentic system evaluation, audit design, calibration facilitation, rubric maintenance, reviewer training, data analysis, Python, SQL
Preferred skills
RLHF experience, red-teaming, trust and safety review, multi-step task execution analysis
Technologies
Python, SQL, spreadsheets, dashboarding tools
Responsibilities
Design audit sampling strategies and scoring standards, re-score reviewer output to identify error patterns, run calibration sessions to resolve disagreements, maintain and revise evaluation rubrics, train and onboard new reviewers, report quality trends and performance evidence to management
Seniority
Senior, hands-on IC