Staff Software Engineer- AI Agent Evaluations
Core
Define and lead the discipline of testing AI agents, evaluating LLM behavior, and ensuring the reliability of agentic systems operating in production.
Role type
Staff Software Engineer (AI Agent Evaluations)
Builds
Evaluation pipelines for LLM outputs, agent behavior, tool use, and multi-turn interactions across development, staging, and production environments.
Domain
AI/ML, Identity Verification, Agentic Systems
Deliverable
production ML models
Required skills
Building and operating production software systems, evaluating/testing LLM-powered features or autonomous agents, backend engineering (Python, Java, Go), designing test infrastructure and CI/CD quality gates, building internal developer tooling, leading cross-team technical initiatives, building eval frameworks for LLM agents, familiarity with agentic frameworks, production monitoring for AI systems, red-teaming or adversarial testing experience.
Preferred skills
Background in identity verification or fraud detection, familiarity with Anthropic's model evaluation methodology, experience with observability tooling (Datadog, OpenTelemetry) for AI workloads, track record of building widely adopted developer platforms.
Technologies
Python, Java, Go, Claude API, Anthropic SDK, BrainTrust, LangChain, LangGraph, CrewAI, Datadog, OpenTelemetry
Responsibilities
Define AI quality standards for evaluating and monitoring AI agents; Build eval infrastructure for LLM outputs and agent behavior; Establish production observability for behavioral drift and failure modes; Design test suites handling non-determinism (red-teaming, golden datasets, LLM-as-judge); Build internal tooling to improve developer experience for AI features; Drive AI-first engineering culture and mentorship; Partner with cross-functional teams to embed quality gates.
Seniority
Staff, hands-on IC with mentorship