Machine Learning Engineer, Evals
Core
Designing and maintaining evaluation infrastructure, benchmarks, and LLM-as-judge systems to assess agent capabilities and model performance.
Role type
Machine Learning Engineer (Evaluation & Benchmarking)
Builds
Evaluation pipelines, judge calibration protocols, extended benchmarks, and failure analysis tooling.
Domain
Artificial Intelligence, Large Language Models (LLMs), Agent Evaluation
Deliverable
production ML models
Required skills
Python, LLM evaluation frameworks, LLM prompting and fine-tuning, Git/CI/CD, Docker, Linux command line, statistical evaluation metrics, failure analysis, benchmark design
Preferred skills
RLVR/RLHF pipeline experience, training data curation, distributed eval orchestration, red teaming, psychometrics
Technologies
Harbor, Nemo Evaluator, GAIA, τ-Bench, SWE-bench, Git, Docker, Linux
Responsibilities
Run full eval pipelines end-to-end and reproduce known results; Build judge calibration protocols; Extend existing benchmarks with new tasks; Run failure analysis on model outputs; Own recurring eval workflows and ship tooling.
Seniority
Mid-level, hands-on IC