Staff Software Engineer, Agent Eval Platform
Core
Build the judgement layer of an agent evaluation platform that scores multi-step agent trajectories in stateful enterprise environments to enable self-correction and optimization.
Role type
Staff Software Engineer (Agent Eval Platform)
Builds
Eval orchestration runtime, agent observability/tracing infrastructure, and stateful simulation environments for enterprise systems.
Domain
Agentic AI, Distributed Systems, Observability, Simulation Engineering
Deliverable
production ML models | product features | infrastructure
Required skills
Distributed systems design, Workflow orchestration, Observability internals, Concurrent/async programming, Data-intensive pipelines, gRPC/protobuf design
Preferred skills
Python, Go, LLM evaluation methodology, Stateful simulation design
Technologies
OpenTelemetry, Python, Go, gRPC, protobuf, Temporal, Airflow, Argo
Responsibilities
Design and build high-concurrency eval orchestration runtimes; Lead implementation of OpenTelemetry-native observability for agent trajectories; Develop stateful simulation environments for enterprise systems; Establish reliability floors and SLOs for the evaluation harness.
Seniority
Staff, hands-on IC with strategic scope