Senior Research Scientist, Model Evaluation
Core
Developing next-generation evaluation methods and infrastructure to measure Large Language Model (LLM) progress and capabilities.
Role type
Senior Research Scientist, Model Evaluation
Builds
Scalable evaluation benchmarks, LLM judges, and data synthesis pipelines for measuring model performance.
Domain
Artificial Intelligence / Large Language Models (LLM)
Deliverable
production ML models
Required skills
LLM evaluation methodology design, LLM-based data synthesis, software engineering, prototype development, data quality review, rigorous measurement frameworks
Preferred skills
Experience building resources to measure LLM capabilities, training LLM judges, improving evaluation efficiency
Technologies
LLMs, evaluation infrastructure, data synthesis pipelines
Responsibilities
Create ambitious new evaluation benchmarks, translate model feedback into trustworthy evaluations, conduct research to advance LLM evaluation methods, build scalable tools for analyzing model performance
Seniority
Senior, hands-on IC