Staff Back End Engineer, Evals - Hazel AI
Core
Architecting and building the evaluation platform for an AI engine in wealth management, ensuring quality, safety, and compliance for financial advice.
Role type
Staff Back End Engineer (AI Evaluation Infrastructure)
Builds
Production observability, monitoring, golden datasets, LLM verification agents, and CI/CD integration for AI quality.
Domain
Wealth management / Financial services / Applied AI
Deliverable
production ML models | infrastructure
Required skills
Evaluation and scoring methodologies for modern AI systems, data curation and golden dataset design, backend integration and API design, observability tooling, SQL and data warehousing, LLM-as-judge frameworks, human-in-the-loop workflow design, regulated industry compliance knowledge.
Preferred skills
Experience with agentic workflows and RAG pipelines, familiarity with evaluation frameworks like Braintrust or Langfuse, background in regulated industries, experience building human-in-the-loop labeling tooling.
Technologies
SQL, dbt, data warehouses, APIs, async pipelines, queues, Anthropic, OpenAI, self-hosted models.
Responsibilities
Design and build the evals platform end-to-end including online scoring and regression suites, build production observability for AI quality metrics, architect data curation pipelines for evaluation datasets, develop LLM verification agents to catch errors, integrate evals into deployment pipelines as first-class gates, partner with SMEs to define quality SLOs and alerting.
Seniority
Staff, hands-on IC with architectural responsibility