Staff Machine Learning Engineer, Agent Eval Platform
Core
Building the judgement layer of an agent evaluation platform to score multi-step agent trajectories in enterprise systems, creating calibrated signals for training and evaluation.
Role type
Staff Machine Learning Engineer (Agent Evaluation & Reward Modeling)
Builds
A calibrated judge system (rubrics, validators, LLM judges) and a process reward model for agentic AI.
Domain
Agentic AI, Enterprise Automation, LLM Evaluation
Deliverable
production ML models
Required skills
Applied ML fundamentals, Python, LLM evaluation and fine-tuning, rubric design, human-in-the-loop calibration, reward modeling, offline/online divergence analysis
Preferred skills
LLM-as-judge design, human annotation program management, search ranking/recsys evaluation, SFT/preference tuning, RLHF/RLAIF, agent trajectory analysis, prompt engineering as an engineering discipline
Technologies
Python, LLMs, SFT, RLHF, RLAIF
Responsibilities
Design shared base judges with per-item rubrics; split validation between deterministic validators and LLM judges; implement confidence-based scoring routing; run calibration loops against human labels; fine-tune small judge models; analyze offline/online divergence; build process reward models for agent optimization.
Seniority
Staff, hands-on IC with strategic ownership