CareerPlanSign in

AI Quality & Evaluation Engineer, AI product

Tokyo💼 Full-time🗓 2026-06-03 → 2026-09-27

Core

Define and automate quality evaluation workflows for LLM-based AI products, establishing feedback loops to drive continuous quality improvements.

Role type

AI Quality & Evaluation Engineer (LLM products)

Builds

Automated quality evaluation workflows and feedback loops for AI products

Domain

AI / Large Language Models (LLM) / Automotive Software

Deliverable

production ML models

Required skills

LLM testing, scenario-based testing, red teaming, test automation, evaluation process design, data aggregation, risk communication

Preferred skills

Natural language processing (NLP), user feedback analysis, log analysis

Technologies

LLM frameworks, automation tools, data pipelines

Responsibilities

Define quality dimensions (accuracy, consistency, safety, fairness, UX); Conduct scenario-based, exploratory, and red teaming tests; Analyze LLM outputs for behavioral trends; Design evaluation processes using user feedback; Standardize and automate evaluation workflows; Collect and aggregate evaluation data; Establish recurring quality checks; Build quality feedback loops with dev/MLOps teams; Document and communicate quality risks

Seniority

Mid-level, hands-on IC

Sourced via lever · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.