Research Lead, Training Insights
Core
Develop strategy and lead execution on measuring and characterizing large language model capabilities across training and deployment lifecycles.
Role type
Senior IC research lead (model evaluation)
Builds
Novel long-horizon evaluation methodologies and frameworks for measuring emerging model capabilities
Domain
Artificial Intelligence / Large Language Models / Reinforcement Learning
Deliverable
production ML models
Required skills
Designing and running evaluations for large language models, leading technical projects or teams, designing experiments and writing code, strategic thinking about measurement, synthesizing information across teams, communicating complex technical findings, results-oriented execution
Preferred skills
Building evaluations for long-horizon or agentic tasks, deep familiarity with Reinforcement Learning training dynamics, published research in machine learning evaluation, experience with safety evaluation frameworks and red teaming, background in psychometrics or experimental psychology, managing or mentoring researchers and engineers
Technologies
Reinforcement Learning, large language models
Responsibilities
Build new novel and long-horizon evaluations, Develop novel measurement approaches for understanding how model capabilities emerge and evolve during RL training, Lead strategic evaluation coverage across the company, Shape the evaluation narrative for model releases, Lead and mentor a small team of researchers and research engineers, Design evaluation frameworks that balance scientific rigor with production training schedules
Seniority
Senior, hands-on IC with leadership responsibilities