Researcher, Evaluations
Core
Lead efforts to evaluate frontier AI models on open-ended, real-world office tasks by curating realistic benchmarks, designing grading rubrics, and assessing model performance quantitatively and qualitatively.
Role type
Researcher, AI Evaluation
Builds
Real-world task benchmarks and evaluation rubrics for frontier AI models
Domain
Artificial Intelligence, Real-world task evaluation
Deliverable
production ML models
Required skills
AI model evaluation, benchmark curation, rubric design, quantitative and qualitative assessment, AI tool usage
Preferred skills
Setting up AI-assisted automated workflows
Responsibilities
Curate realistic task suites for AI benchmarking, design grading rubrics for AI performance, run newly-released models through evaluation suites, assess model performance quantitatively and qualitatively
Sourced via lever · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.