CareerPlanSign in

Machine Learning Engineer, Evals

Americas (US time zones)💼 Full-time🗓 2026-08-13 → 2026-09-25

Core

Designing and maintaining evaluation infrastructure, benchmarks, and LLM-as-judge systems to assess agent capabilities and model performance.

Role type

Machine Learning Engineer (Evaluation & Benchmarking)

Builds

Evaluation pipelines, judge calibration protocols, extended benchmarks, and failure analysis tooling.

Domain

Artificial Intelligence, Large Language Models (LLMs), Agent Evaluation

Deliverable

production ML models

Required skills

Python, LLM evaluation frameworks, LLM prompting and fine-tuning, Git/CI/CD, Docker, Linux command line, statistical evaluation metrics, failure analysis, benchmark design

Preferred skills

RLVR/RLHF pipeline experience, training data curation, distributed eval orchestration, red teaming, psychometrics

Technologies

Harbor, Nemo Evaluator, GAIA, τ-Bench, SWE-bench, Git, Docker, Linux

Responsibilities

Run full eval pipelines end-to-end and reproduce known results; Build judge calibration protocols; Extend existing benchmarks with new tasks; Run failure analysis on model outputs; Own recurring eval workflows and ship tooling.

Seniority

Mid-level, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.