Tech Lead, Agent Eval Platform
Core
Build the judgement layer of an agent evaluation platform, including rubrics, judges, calibration against human labels, and methodology to score multi-step agent trajectories in stateful enterprise environments.
Role type
Tech Lead, Applied ML Infrastructure & Evaluation Systems
Builds
A scalable runtime for multi-turn agent scenarios, an OpenTelemetry-native observability stack for agent trajectories, and a stateful simulation environment for enterprise systems.
Domain
Agentic AI, Enterprise Automation, Distributed Systems, Observability
Deliverable
production ML models | infrastructure
Required skills
Distributed systems design, Workflow orchestration (DAGs, scheduling), Observability internals (OpenTelemetry), Concurrent/async programming, Data-intensive pipelines, gRPC/protobuf design
Preferred skills
Python and Go, High-concurrency system operation, Non-deterministic system measurement, Ambiguity navigation
Technologies
Python, Go, OpenTelemetry, gRPC, protobuf, Temporal, Airflow, Argo
Responsibilities
Design and operate high-concurrency execution runtimes for agent evals; Lead the implementation of a unified span data model for agent tracing; Build stateful simulation environments with programmatic setup/teardown; Establish reliability floors and SLOs for the evaluation harness; Mentor engineers on end-to-end delivery of novel evaluation methodologies.
Seniority
Senior, hands-on IC with leadership responsibilities