CareerPlanGet AI match score →

Data Scientist 5 - AI Evals

USA - Remote💼 Full-time💰 $372,000–$372,000🗓 2026-06-01 → 2026-07-16

Core

Architect systems and frameworks to measure, validate, and optimize GenAI systems in production for Netflix Games, bridging model capabilities with end-user experience.

Role type

Senior Data Scientist (AI Evals)

Builds

Evaluation pipelines, golden datasets, test suites, and safety evaluators for GenAI-powered games and internal agentic tools.

Domain

Interactive entertainment / Generative AI

Deliverable

production ML models

Required skills

experimental design, causal inference, A/B testing, uncertainty quantification, LLM-as-a-Judge frameworks, human-in-the-loop grading, simulation-based testing, prompt engineering, RAG Evals, agentic Evals, agent architectures, long-horizon reasoning evaluation, red-teaming protocols, toxicity detection, hallucination mitigation.

Preferred skills

defining core user experience metrics in gaming, game development team collaboration, MLOps best practices.

Technologies

OpenAI evaluation suites, Anthropic evaluation suites

Responsibilities

Partner with GenAI research team to ensure product graduation from R&D to production; Build and operate robust evaluation pipelines on production-stage GenAI experiences; Curate high-quality golden datasets, test suites, adversarial challenge sets, and synthetic testbeds; Design experiments to understand trade-offs between technical attributes and end-user experience quality; Measure the coherence, fluency, relevance, and joy value of AI-powered game features; Design red-teaming protocols and safety evaluators to detect and mitigate toxicity, hallucinations, jailbreaks, and out-of-character behavior; Guide Evals for internal agentic tools.

Seniority

Senior, hands-on IC

Sourced via eightfold · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Netflix ↗