CareerPlanSign in

Research Engineer, Benchmarks

San Francisco🌐 Remote💼 Full-time🗓 2026-09-22 → 2026-09-29

Core

Design, implement, and maintain high-quality benchmarks for evaluating frontier AI agents on domain-specific tasks.

Role type

Research Engineer (AI Benchmarks)

Builds

Technically rigorous and credible benchmarks for frontier labs

Domain

Artificial Intelligence / Frontier AI Agents / Evaluation

Deliverable

production ML models | product features

Required skills

Python, Docker, Linux, benchmark design, task definition, metrics development, infrastructure building, documentation

Preferred skills

First-principles reasoning, edge case identification, unstructured problem solving, independent work in fast-paced environments

Technologies

Python, Docker, Linux

Responsibilities

Design and implement internal agent benchmarks; collaborate with subject-matter experts to define domain-specific tasks; build infrastructure to run models against benchmarks; develop metrics to analyze benchmark difficulty and failure modes; validate benchmark correlation with real-world needs; write documentation and reports for technical audiences

Seniority

Individual Contributor, early-stage startup

Sourced via ashby · Listed on CareerPlan, which tracks 847,000+ jobs from 20+ sources.