CareerPlanSign in

Sr. Evaluation Engineer

San Francisco, CA💼 Full-time🗓 2026-08-19 → 2026-09-26

Core

Design and build evaluation systems, pipelines, and frameworks to guide the development, testing, and release of Edwin AI, an AI-powered observability and incident intelligence platform.

Role type

Senior AI Evaluation Engineer

Builds

Production-grade evaluation pipelines, golden datasets, automated graders, and regression frameworks for AI agents and retrieval systems.

Domain

AI/ML Engineering, Observability, IT Operations

Deliverable

production ML models

Required skills

Python engineering, AI evaluation frameworks (LangSmith, Arize, DeepEval, etc.), LLMs and agents, retrieval-augmented generation, prompt engineering, tool calling, regression testing, CI/CD integration, behavioral drift detection, LLM-as-a-judge techniques, test case generation, rubric design.

Preferred skills

Experience with multi-step/multi-agent AI systems, human-in-the-loop evaluation, incident investigation workflows, ITSM integrations.

Technologies

Python, LangSmith, Arize Phoenix, Braintrust, DeepEval, Ragas, TruLens, OpenAI Evals, MLflow, CI/CD tools.

Responsibilities

Define quality metrics for incident diagnostics and root-cause analysis; build offline/online evaluation pipelines; create and maintain golden datasets and regression suites; design step-level and trajectory-level evaluations for multi-agent workflows; calibrate LLM-based graders against human judgment; monitor AI quality and behavioral drift in production; establish evaluation-driven development practices.

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.