Staff Back End Engineer, Evals - Hazel AI
Core
Architect and build the evaluation platform for an AI wealth management engine, ensuring quality, safety, and compliance for financial advisors and clients.
Role type
Staff Back End Engineer (AI Evaluation Infrastructure)
Builds
Production observability, monitoring, golden datasets, LLM verification agents, and CI/CD integration for AI quality.
Domain
Wealth management / Financial services / AI Safety & Evaluation
Deliverable
production ML models | infrastructure
Required skills
Evaluation infrastructure design, RAG evaluation, golden dataset curation, LLM-as-judge frameworks, data engineering (SQL, dbt), backend integration, observability tooling
Preferred skills
Agentic workflow evaluation, human-in-the-loop labeling, regulated industry experience, wealth management domain knowledge
Technologies
Anthropic, OpenAI, self-hosted models, Braintrust, Langfuse
Responsibilities
Design and build the evals platform end-to-end including online scoring and regression suites; Build production observability for AI quality metrics; Architect data curation pipelines for evaluation datasets; Develop LLM verification agents to catch hallucinations and compliance violations; Integrate evals into deployment pipelines as a first-class gate; Define quality SLOs and build alerting for production regressions.
Seniority
Staff, hands-on IC with architectural leadership