CareerPlanGet AI match score →

Senior Research Scientist, Model Evaluation

London, England, UK🌐 Remote💼 Full-time🗓 2026-06-15 → 2026-07-29

Core

Developing next-generation evaluation methods and infrastructure to measure Large Language Model (LLM) progress and capabilities.

Role type

Senior Research Scientist, Model Evaluation

Builds

Scalable evaluation benchmarks, LLM judges, and data synthesis pipelines for measuring model performance.

Domain

Artificial Intelligence / Large Language Models (LLM)

Deliverable

production ML models

Required skills

LLM evaluation methodology design, LLM-based data synthesis, software engineering, prototype development, data quality review, rigorous measurement alignment

Preferred skills

Experience measuring LLM capabilities, building resources to measure model boundaries

Technologies

LLMs, evaluation frameworks, data synthesis pipelines

Responsibilities

Create ambitious new evaluation benchmarks, translate model feedback into trustworthy evaluations, conduct research to advance LLM evaluation methods, build scalable tools for analyzing model performance

Seniority

Senior, hands-on IC

Sourced via linkedin · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on LinkedIn ↗