Staff Software Development Test Engineer - AI Evaluation
Core
Building the evaluation platform and frameworks that enable ML teams to measure, trust, and improve the quality of AI outputs across the organization.
Role type
Staff Software Development Test Engineer (AI Evaluation)
Builds
Shared AI evaluation infrastructure, automated scoring pipelines, and quality monitoring systems for ML agents.
Domain
Automotive retail technology / AI/ML evaluation
Deliverable
production ML models
Required skills
Python, ML/LLM evaluation, benchmark design, LLM-as-judge, dataset curation, statistical analysis, CI/CD integration, dashboarding
Preferred skills
Eval frameworks (Ragas, DeepEval, LangSmith), online evaluation/guardrails, responsible AI/safety evaluation, experiment tracking (MLflow, W&B)
Technologies
Python, LLMs, CI/CD pipelines, MLflow, W&B, Ragas, DeepEval, LangSmith, TruLens, Promptfoo, HELM
Responsibilities
Design and own AI evaluation infrastructure; create and maintain evaluation datasets and ground-truth sets; define quality metrics for AI outputs; build automated scoring pipelines; validate user intents and measure response accuracy; identify hallucinations and edge cases; establish evaluation gates in CI/CD; develop dashboards for AI quality visibility; use AI/LLMs to scale evaluation via automated judges and synthetic data.
Seniority
Staff, hands-on IC with strategic platform ownership