Research Engineer, Benchmarks
Core
Design, implement, and maintain high-quality benchmarks for evaluating frontier AI agents on domain-specific tasks.
Role type
Research Engineer (AI Benchmarks)
Builds
Technically rigorous and credible benchmarks for frontier labs
Domain
Artificial Intelligence / Frontier AI Agents / Evaluation
Deliverable
production ML models | product features
Required skills
Python, Docker, Linux, benchmark design, task definition, metrics development, infrastructure building, documentation
Preferred skills
First-principles reasoning, edge case identification, unstructured problem solving, independent work in fast-paced environments
Technologies
Python, Docker, Linux
Responsibilities
Design and implement internal agent benchmarks; collaborate with subject-matter experts to define domain-specific tasks; build infrastructure to run models against benchmarks; develop metrics to analyze benchmark difficulty and failure modes; validate benchmark correlation with real-world needs; write documentation and reports for technical audiences
Seniority
Individual Contributor, early-stage startup
