Agentic AI Evaluation Engineer
Core
Designing and developing evaluation pipelines, metrics, and test cases to validate the behavior, accuracy, and reliability of AI agents before release.
Role type
Agentic AI Evaluation Engineer
Builds
Automated and human-in-the-loop evaluation systems, evaluation pipelines, and CI/CD integrations for AI agents
Domain
Artificial Intelligence / Large Language Models (LLMs)
Deliverable
production ML models
Required skills
AI Agents, Benchmarking, CI/CD, Evaluation Metrics, Large Language Models (LLMs), Machine Learning (ML), Curious Mindset
Preferred skills
Customer support AI or chatbot platforms, Responsible AI (bias, fairness, hallucination mitigation)
Technologies
CI/CD pipelines, LLMs
Responsibilities
Design and develop agent evaluation pipelines across development, staging, and production environments; Define and standardize evaluation metrics and benchmarks for conversational AI quality; Build automated and human-in-the-loop evaluation systems; Manage and curate evaluation datasets, test sets, and annotation workflows; Enable continuous evaluation and monitoring of agents in production; Integrate evaluation into CI/CD pipelines; Conduct experiments, A/B testing, and case studies to drive improvements in agent quality; Create technical documentation and drive best practices across teams; Mentor junior engineers and contribute to team growth
Seniority
Mid-Senior (5-7 years experience)