Lead AI Engineer 4C
Core
Design and execute comprehensive evaluation frameworks, benchmarks, and success metrics for large-scale AI models and agentic AI systems.
Role type
Senior AI evaluation engineer (model testing & safety)
Builds
Scalable evaluation processes, test datasets, and risk-mitigation plans for AI products
Domain
Artificial Intelligence / Machine Learning / Responsible AI
Deliverable
production ML models
Required skills
AI/ML evaluation, model testing, human-in-the-loop evaluation design, Python, data analysis, experiment tracking, prompt engineering, AI/ML Ops
Preferred skills
Advanced analytics, prompt design, safety strategy development
Technologies
Python, experiment-tracking frameworks
Responsibilities
Develop evaluation methodologies for performance, safety, robustness, and fairness; lead large-scale human evaluations and create test datasets; analyze results to identify failure patterns and translate findings into actionable recommendations; partner with researchers and product teams to align evaluation goals; ensure compliance with safety standards and regulatory expectations.
Seniority
Senior, hands-on IC