CareerPlanGet AI match score →

Staff Back End Engineer, Evals - Hazel AI

San Francisco, CA💼 Full-time💰 $275,000–$275,000🗓 2026-05-27 → 2026-07-31

Core

Architect and build the evaluation platform for an AI wealth management engine, ensuring quality, safety, and compliance for financial advisors and clients.

Role type

Staff Back End Engineer (AI Evaluation Infrastructure)

Builds

Production observability, monitoring, golden datasets, LLM verification agents, and CI/CD integration for AI quality.

Domain

Wealth management / Financial services / AI Safety & Evaluation

Deliverable

production ML models | infrastructure

Required skills

Evaluation infrastructure design, RAG evaluation, golden dataset curation, LLM-as-judge frameworks, data engineering (SQL, dbt), backend integration, observability tooling

Preferred skills

Agentic workflow evaluation, human-in-the-loop labeling, regulated industry experience, wealth management domain knowledge

Technologies

Anthropic, OpenAI, self-hosted models, Braintrust, Langfuse

Responsibilities

Design and build the evals platform end-to-end including online scoring and regression suites; Build production observability for AI quality metrics; Architect data curation pipelines for evaluation datasets; Develop LLM verification agents to catch hallucinations and compliance violations; Integrate evals into deployment pipelines as a first-class gate; Define quality SLOs and build alerting for production regressions.

Seniority

Staff, hands-on IC with architectural leadership

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗