CareerPlanSign in
💼 Full-time🗓 2026-06-25 → 2026-09-24

Core

Building the unified evaluation infrastructure, automated pipelines, and production feedback loops to ensure AI agents perform reliably at enterprise scale for audit and advisory workflows.

Role type

Senior AI Engineer (Quality & Evaluation Infrastructure)

Builds

Unified evaluation platform, automated evaluation pipelines, observability systems, and production feedback loops for agentic systems.

Domain

AI Agents, Audit & Advisory, Enterprise Software

Deliverable

production ML models | infrastructure

Required skills

TypeScript, React, Python, Postgres, LLM orchestration, agent design, evaluation framework design, observability/tracing, distributed systems, cloud infrastructure

Preferred skills

AI-native engineering instincts, data-driven decision making, product judgment, rapid prototyping

Technologies

LangSmith, LangGraph, TypeScript, React, Python, Postgres

Responsibilities

Design and build a unified evaluation platform as the single source of truth for agentic systems; Build observability systems to surface agent behavior and failure modes; Implement automated pipelines to evaluate new models against critical workflows within hours; Design guardrails and monitoring systems to catch quality regressions; Integrate and orchestrate LLMs, tools, and retrieval systems into reliable agent experiences; Define and document evaluation standards and best practices for the engineering organization.

Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.