Senior AI Engineer, Quality
Core
Building the unified evaluation infrastructure, automated pipelines, and production feedback loops to ensure AI agents perform reliably at enterprise scale for audit and advisory workflows.
Role type
Senior AI Engineer (Quality & Evaluation Infrastructure)
Builds
Unified evaluation platform, observability systems, automated model evaluation pipelines, and guardrails for agentic systems.
Domain
AI/ML Engineering, Observability, Quality Assurance, Audit & Advisory
Deliverable
production ML models | infrastructure
Required skills
TypeScript, Python, Postgres, LLM orchestration, RAG architectures, vector databases, evaluation framework design, observability/tracing, LangSmith/LangGraph integration, prompt engineering, data-driven decision making.
Preferred skills
Experience shipping production LLM features, building evaluation harnesses for complex agentic systems, designing comparison frameworks for model effectiveness/latency/cost, working with SMEs to curate production traces.
Responsibilities
Design and build a unified evaluation platform as the single source of truth; build observability systems to surface agent behavior and failure modes; implement automated pipelines to evaluate new models within hours; design guardrails and monitoring to catch quality regressions; integrate and orchestrate LLMs, tools, and retrieval systems; define evaluation standards and advocate for evaluation-driven development.
Seniority
Senior, hands-on IC with ownership of large product areas.