Research Engineer, Benchmarks
Core
Designing and owning high-quality benchmarks to evaluate frontier AI agents on realistic, domain-specific workflows.
Role type
Research Engineer (AI Benchmarks & Evaluation)
Builds
Reliable infrastructure for running models and agents against benchmark tasks at scale
Domain
Artificial Intelligence / Machine Learning / Evaluation Infrastructure
Deliverable
production ML models
Required skills
Python, Docker, Linux, benchmark design, statistical analysis, workflow modeling, technical documentation
Preferred skills
reinforcement learning pipelines, data generation, RL agent evaluation, published work on AI benchmarking
Technologies
Python, Docker, Linux
Responsibilities
Design and implement internal benchmarks for evaluating frontier agents; Partner with subject-matter experts to define realistic workflows and evaluation criteria; Build and operate infrastructure to run models and agents at scale; Develop metrics and statistical analyses for benchmark difficulty and reliability; Validate benchmark performance against real-world evaluations; Write technical documentation and benchmark reports
Seniority
Mid-level, hands-on IC