Senior Software Engineer, AI Evals
Core
Building evaluation infrastructure to measure the accuracy, reliability, and real-world performance of AI systems, including debugging agents and AI-powered features.
Role type
Senior IC software engineer specializing in AI evaluation infrastructure
Builds
Evaluation frameworks, datasets, benchmarks, test harnesses, and metrics pipelines for AI models and agentic workflows
Domain
AI/ML engineering, software reliability, and application monitoring
Deliverable
production ML models | infrastructure
Required skills
Python, TypeScript, building testing/evaluation/data infrastructure, working with structured/unstructured datasets, designing benchmarks and test harnesses, defining measurable criteria for AI behavior
Preferred skills
evaluating LLMs, agentic systems, AI-assisted developer tools, offline/online evaluation techniques, regression testing for models or prompts
Technologies
Python, TypeScript
Responsibilities
Design and build robust evaluation frameworks; Create and curate high-quality datasets and golden test cases; Build automated test harnesses and metrics pipelines; Partner with applied AI engineers and product leaders to define measurable criteria; Own the evaluation lifecycle from experimentation to production monitoring
Seniority
Senior, hands-on IC