Data Scientist 5 - AI Evals
Core
Architect systems and frameworks to measure, validate, and optimize GenAI systems in production for Netflix Games, bridging model capabilities with end-user experience.
Role type
Senior Data Scientist (AI Evals)
Builds
Evaluation pipelines, golden datasets, test suites, and safety evaluators for GenAI-powered games and internal agentic tools.
Domain
Interactive entertainment / Generative AI
Deliverable
production ML models
Required skills
experimental design, causal inference, A/B testing, uncertainty quantification, LLM-as-a-Judge frameworks, human-in-the-loop grading, simulation-based testing, prompt engineering, RAG Evals, agentic Evals, agent architectures, long-horizon reasoning evaluation, red-teaming protocols, toxicity detection, hallucination mitigation.
Preferred skills
defining core user experience metrics in gaming, game development team collaboration, MLOps best practices.
Technologies
OpenAI evaluation suites, Anthropic evaluation suites
Responsibilities
Partner with GenAI research team to ensure product graduation from R&D to production; Build and operate robust evaluation pipelines on production-stage GenAI experiences; Curate high-quality golden datasets, test suites, adversarial challenge sets, and synthetic testbeds; Design experiments to understand trade-offs between technical attributes and end-user experience quality; Measure the coherence, fluency, relevance, and joy value of AI-powered game features; Design red-teaming protocols and safety evaluators to detect and mitigate toxicity, hallucinations, jailbreaks, and out-of-character behavior; Guide Evals for internal agentic tools.
Seniority
Senior, hands-on IC