Senior Research Scientist, Model Evaluation
Core
Developing next-generation evaluation methods and infrastructure to measure Large Language Model (LLM) progress and capabilities.
Role type
Senior Research Scientist, Model Evaluation
Builds
Scalable evaluation benchmarks, LLM judges, and data synthesis pipelines for measuring model performance.
Domain
Artificial Intelligence / Large Language Models (LLM)
Deliverable
production ML models
Required skills
LLM evaluation methodology design, LLM-based data synthesis, software engineering, prototype development, data quality review, rigorous measurement alignment
Preferred skills
Experience measuring LLM capabilities, building resources to measure model boundaries
Technologies
LLMs, evaluation frameworks, data synthesis pipelines
Responsibilities
Create ambitious new evaluation benchmarks, translate model feedback into trustworthy evaluations, conduct research to advance LLM evaluation methods, build scalable tools for analyzing model performance
Seniority
Senior, hands-on IC