Member of Technical Staff, Enterprise Evals Platform
Core
Build the evaluation platform, verifiers, and grading infrastructure for enterprise AI agents to ensure reliability and economic viability.
Role type
Senior IC platform engineer specializing in AI agent evaluation systems
Builds
Scalable evaluation platform, verifiers, offline environments, and task suites for agent measurement
Domain
AI/ML, Agent Systems, Enterprise Software
Deliverable
production ML models | infrastructure
Required skills
Agent engineering and evaluation, LLM benchmark construction, rubric design, software engineering fundamentals, loss analysis, optimization loops, rollout gate ownership
Preferred skills
Harbor environments, RL environments
Technologies
LLMs, agent runtimes, harnesses, terminal-bench, tau-bench, APEX
Responsibilities
Define golden sets by decomposing real tasks and encoding expert quality bars; Build verifiers over agent trajectories and outputs; Build the eval platform running offline environments and grading at scale; Run loss analysis over production trajectories; Run optimization loops across models and prompts; Own rollout gates for agent changes
Seniority
Senior, hands-on IC