Software Engineering Director, Agentic Evaluations
Core
Lead the delivery of credible evaluations for AI agents offered by software vendors by building and enhancing the core evaluation system.
Role type
Senior IC Software Engineering Director (Agentic Evaluations)
Builds
Full-stack agent evaluation system (admin, API, workflow dispatch) and proprietary benchmarks
Domain
B2B Software Marketplace / AI Agent Evaluation
Deliverable
production ML models | product features
Required skills
Backend development (Python, Java/Kotlin, Typescript/Javascript, Go), Agent evaluation design, LLM-as-a-judge applications, Coding agent harnesses (Claude Code, Codex, Opencode, Pi), System architecture, Team leadership
Preferred skills
Agent tool use (MCP servers), Benchmark frameworks (STATE-Bench, tau2-bench), Browser automation (Playwright, Chrome DevTools MCP), Durable execution frameworks (Temporal, DBOS, Cloudflare/Vercel Workflows), Agent sandboxing (AWS E2B, Daytona)
Technologies
FastAPI, Node.js, Python, Java/Kotlin, Typescript/Javascript, Go, OpenAI, Anthropic, Google, Temporal, DBOS, Cloudflare, Vercel, AWS E2B, Daytona, Playwright, Chrome DevTools MCP
Responsibilities
Scope feasibility for vendor agent evaluations, own full stack of evaluation system, design generalized evaluation primitives, distill repeatable onboarding processes, track emerging agent evaluation frameworks, mentor engineers and socialize eval usage across products
Seniority
Director, hands-on IC with management