Senior Machine Learning Engineer, Agent Eval Platform
Core
Building the judgement layer of an agent evaluation platform to score multi-step agent trajectories in enterprise systems, creating calibrated signals for training and optimization.
Role type
Senior IC machine learning engineer (agent evaluation & reward modeling)
Builds
A calibrated judge artifact, process reward models, and a simulated world substrate for agent optimization.
Domain
Agentic AI, enterprise automation, LLM evaluation
Deliverable
production ML models
Required skills
applied ML fundamentals, Python, LLM evaluation and fine-tuning, rubric design, human annotation program management, offline/online divergence analysis, step-level fault attribution
Preferred skills
LLM-as-judge design, search ranking/recsys evaluation, reward modeling (RLHF/RLAIF), agent trajectory analysis, prompt engineering as an engineering discipline
Technologies
Python, LLMs, simulation environments
Responsibilities
Design shared base judges with per-item rubrics; split validation between deterministic validators and LLM judges; implement scoring with confidence reporting; run calibration loops against human labels; fine-tune small judge models; guard against correlated blind spots; build process reward models for agent optimization; maintain versioned scenarios and simulated worlds.
Seniority
Senior, hands-on IC