Manager, AI Evaluation
Core
Lead the assessment, testing, and governance of generative AI solutions to ensure they are safe, effective, and fit for deployment across the AI Hub.
Role type
Manager, AI Evaluation
Builds
Evaluation frameworks, testing methodologies, and governance processes for generative AI services
Domain
Higher Education / Generative AI & LLMs
Deliverable
production ML models
Required skills
Generative AI expertise, Large Language Models (LLMs), agentic AI frameworks, data science, machine learning, natural language processing, evaluation framework design, regression testing, risk identification, stakeholder engagement
Preferred skills
Postgraduate qualifications in relevant field, experience in complex organizational environments, Agile methodology
Technologies
LLMs, agentic AI frameworks
Responsibilities
Lead development of AI evaluation frameworks and standards; design and govern evaluation datasets and metrics; oversee pre-release testing and governance approvals; manage ongoing monitoring and performance evaluation of production AI services; partner with technical teams and governance bodies; provide leadership and capability development in AI evaluation
Seniority
Manager, hands-on leadership