AI Quality & Evaluation Engineer, AI product
Core
Define and automate quality evaluation workflows for LLM-based AI products, establishing feedback loops to drive continuous quality improvements.
Role type
AI Quality & Evaluation Engineer (LLM products)
Builds
Automated quality evaluation workflows and feedback loops for AI products
Domain
AI / Large Language Models (LLM) / Automotive Software
Deliverable
production ML models
Required skills
LLM testing, scenario-based testing, red teaming, test automation, evaluation process design, data aggregation, risk communication
Preferred skills
Natural language processing (NLP), user feedback analysis, log analysis
Technologies
LLM frameworks, automation tools, data pipelines
Responsibilities
Define quality dimensions (accuracy, consistency, safety, fairness, UX); Conduct scenario-based, exploratory, and red teaming tests; Analyze LLM outputs for behavioral trends; Design evaluation processes using user feedback; Standardize and automate evaluation workflows; Collect and aggregate evaluation data; Establish recurring quality checks; Build quality feedback loops with dev/MLOps teams; Document and communicate quality risks
Seniority
Mid-level, hands-on IC