CareerPlanGet AI match score →

Research Engineer, Model Evaluations

San Francisco, CA💼 Full-time💰 $320,000–$320,000🗓 2026-04-28 → 2026-07-31

Core

Build evaluation metrics and infrastructure to measure Claude's capabilities (reasoning, safety, agentic behavior) and ensure reliable, defensible metrics for researchers and leadership.

Role type

Research Engineer (Model Evaluations)

Builds

Distributed evaluation execution platform, dashboards for model health monitoring, and evaluation tooling/libraries.

Domain

AI Safety / Large Language Model Evaluation

Deliverable

production ML models | dashboards & analysis

Required skills

Python programming, distributed systems, data pipelines, production support/on-call, technical communication

Preferred skills

LLM prompting/sampling/scaffolding, data visualization, evaluation metrics design, observability, statistics/experimental design, dataset curation, ML training infrastructure

Technologies

Python, distributed systems, dashboards, observability tools, experiment tracking

Responsibilities

Design and run evaluations of Claude's capabilities; build and harden distributed eval execution platform; own dashboards for model health monitoring; debug anomalous eval results mid-training; improve researcher tooling and workflows; partner with research teams on capability lifecycles; run experiments on prompting/sampling effects; communicate results to stakeholders.

Seniority

Mid-to-Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗