Director, Evaluations
Core
Define and lead world-class evaluation strategy, infrastructure, and teams for safe-by-design AI systems (Scientist AI) and frontier LLMs, ensuring independent verification of capability and safety claims.
Role type
Director of Evaluations (Leadership)
Builds
Independent evaluation infrastructure, datasets, benchmarks, red-teaming programs, and automated pipelines for AI safety and capability assessment.
Domain
AI Safety / Machine Learning / Frontier LLMs
Deliverable
production ML models | research | infrastructure
Required skills
Strategic roadmap definition, team building and scaling, independent evaluation design, dataset/benchmark creation, red-teaming program leadership, automated pipeline construction, external stakeholder communication
Preferred skills
Third-party red-teaming partner management, open-source dataset release, AI safety policy standards familiarity, external safety institute coordination
Technologies
LLMs, frontier ML systems, automated evaluation tooling, interactive environments
Responsibilities
Define evaluation strategy and roadmap; build and scale the Evaluations Team; operate independently of research streams to avoid conflicts of interest; design novel benchmarks for capabilities and safety; oversee evaluation of Scientist AI as a guardrail; lead automated and manual red-teaming programs; construct internal tooling for scale; support research/product streams with evaluation requirements; own public communication of evaluation results; represent LawZero externally on AI safety measurement.
Seniority
Director, strategic leadership & team building