AI Quality Engineer
Core
Build and maintain evaluation systems and test infrastructure to ensure confidence in non-deterministic AI output for due diligence and risk intelligence.
Role type
Senior AI Quality Engineer (LLM Evaluation & Test Infrastructure)
Builds
Automated evaluation harnesses, versioned golden datasets, and CI/CD pipelines for AI investigative output.
Domain
AI-driven due diligence, risk intelligence, and financial services.
Deliverable
production ML models
Required skills
Python, Playwright, pytest, CI/CD, LLM evaluation (rubrics, LLM-as-judge), automated testing, data analysis, root cause analysis, cross-functional communication.
Preferred skills
Experience shipping LLM evals, experience with non-deterministic output testing, ability to simplify complex technical concepts.
Technologies
Python, Playwright, pytest, CI/CD, Claude Code.
Responsibilities
Enhance and run layered eval harnesses including automated checks and human review; maintain versioned golden datasets; track signals like groundedness and hallucination rates; run and improve tiered CI models; fix flaky tests and wobbling metrics.
Seniority
Senior, hands-on IC