Engagement Manager, Agentic AI Workflow Evaluations
Core
Lead a dedicated onsite team of 11 reviewers and 1 QA lead to evaluate complex, real-world agentic AI workflows for a frontier AI customer, ensuring safety, user intent respect, and task completion accuracy.
Role type
Senior delivery manager (agentic AI evaluation)
Builds
Evaluation frameworks and human expertise for trusted AI systems
Domain
Artificial Intelligence / Agentic AI / Trust and Safety
Deliverable
production ML models
Required skills
Delivery team management (10-20 people), customer-facing engagement ownership, AI/ML evaluation (annotation, red-teaming, RLHF, agent evaluation), rubric auditing, escalation management, forecasting staffing and cost
Preferred skills
Experience with frontier AI customers, technical depth to audit reviewer output and defend scoring decisions
Technologies
Agentic AI workflows, isolated test environments
Responsibilities
Own end-to-end delivery for throughput, quality, and capacity; manage and develop a team of eleven including hiring and performance management; serve as primary onsite point of contact for customer program and technical leads; audit reviewer output by sampling trajectories and checking rubric application; partner with QA lead on calibration cycles and rubric refinement; own escalation path for safety-relevant findings; identify expansion opportunities within the account; forecast staffing and cost against commercial model; maintain information security and privacy practices
Seniority
Senior, hands-on IC with management responsibilities
