CareerPlanGet AI match score →

Staff Back End Engineer, Evals - Hazel AI

San Francisco, CA, US💼 Full-time💰 $275,000–$325,000🗓 2026-05-30 → 2026-07-28

Core

Architecting and building the evaluation platform for an AI engine in wealth management, ensuring quality, safety, and compliance for financial advice.

Role type

Staff Back End Engineer (AI Evaluation Infrastructure)

Builds

Production observability, monitoring, golden datasets, LLM verification agents, and CI/CD integration for AI quality.

Domain

Wealth management / Financial services / Applied AI

Deliverable

production ML models | infrastructure

Required skills

Evaluation and scoring methodologies for modern AI systems, data curation and golden dataset design, backend integration and API design, observability tooling, SQL and data warehousing, LLM-as-judge frameworks, human-in-the-loop workflow design, regulated industry compliance knowledge.

Preferred skills

Experience with agentic workflows and RAG pipelines, familiarity with evaluation frameworks like Braintrust or Langfuse, background in regulated industries, experience building human-in-the-loop labeling tooling.

Technologies

SQL, dbt, data warehouses, APIs, async pipelines, queues, Anthropic, OpenAI, self-hosted models.

Responsibilities

Design and build the evals platform end-to-end including online scoring and regression suites, build production observability for AI quality metrics, architect data curation pipelines for evaluation datasets, develop LLM verification agents to catch errors, integrate evals into deployment pipelines as first-class gates, partner with SMEs to define quality SLOs and alerting.

Seniority

Staff, hands-on IC with architectural responsibility

Sourced via efinancialcareers · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on eFinancialCareers ↗