Sr. Evaluation Engineer
Core
Design and build evaluation systems, pipelines, and frameworks to guide the development, testing, and release of Edwin AI, an AI-powered observability and incident intelligence platform.
Role type
Senior AI Evaluation Engineer
Builds
Production-grade evaluation pipelines, golden datasets, automated graders, and regression frameworks for AI agents and retrieval systems.
Domain
AI/ML Engineering, Observability, IT Operations
Deliverable
production ML models
Required skills
Python engineering, AI evaluation frameworks (LangSmith, Arize, DeepEval, etc.), LLMs and agents, retrieval-augmented generation, prompt engineering, tool calling, regression testing, CI/CD integration, behavioral drift detection, LLM-as-a-judge techniques, test case generation, rubric design.
Preferred skills
Experience with multi-step/multi-agent AI systems, human-in-the-loop evaluation, incident investigation workflows, ITSM integrations.
Technologies
Python, LangSmith, Arize Phoenix, Braintrust, DeepEval, Ragas, TruLens, OpenAI Evals, MLflow, CI/CD tools.
Responsibilities
Define quality metrics for incident diagnostics and root-cause analysis; build offline/online evaluation pipelines; create and maintain golden datasets and regression suites; design step-level and trajectory-level evaluations for multi-agent workflows; calibrate LLM-based graders against human judgment; monitor AI quality and behavioral drift in production; establish evaluation-driven development practices.
Seniority
Senior, hands-on IC