CareerPlanGet AI match score →

Senior Research Scientist, Model Evaluation

Toronto🌐 Remote💼 Full-time🗓 2025-10-22 → 2026-07-31

Core

Developing next-generation evaluation methods and infrastructure to measure Large Language Model (LLM) progress and capabilities.

Role type

Senior Research Scientist, Model Evaluation

Builds

Scalable evaluation benchmarks, LLM judges, and data synthesis pipelines for measuring model performance.

Domain

Artificial Intelligence / Large Language Models (LLM)

Deliverable

production ML models

Required skills

LLM evaluation methodology design, LLM-based data synthesis, software engineering, prototype development, data quality review, rigorous measurement frameworks

Preferred skills

Experience building resources to measure LLM capabilities, training LLM judges, improving evaluation efficiency

Technologies

LLMs, evaluation infrastructure, data synthesis pipelines

Responsibilities

Create ambitious new evaluation benchmarks, translate model feedback into trustworthy evaluations, conduct research to advance LLM evaluation methods, build scalable tools for analyzing model performance

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗