CareerPlanGet AI match score →

Senior Research Engineer, Model Evaluation

Toronto🌐 Remote💼 Full-time🗓 2025-07-08 → 2026-07-31

Core

Developing next-generation evaluation methods and scalable infrastructure to measure the performance and capabilities of frontier Large Language Models (LLMs).

Role type

Senior Research Engineer, Model Evaluation

Builds

Evaluation benchmarks, datasets, environments, and scalable analysis tools for LLMs.

Domain

Artificial Intelligence / Large Language Models (LLMs)

Deliverable

production ML models | research

Required skills

LLM evaluation method development, dataset creation, simulator/environment building, software engineering, performance analysis

Preferred skills

Publications at top-tier conferences, experience with popular benchmarks, deep LLM usage experience

Technologies

LLMs, evaluation frameworks, data pipelines

Responsibilities

Develop evaluation benchmarks, datasets, and environments for measuring bleeding-edge model capabilities; Conduct research to push the state-of-the-art in LLM evaluation methods, including training LLM judges; Build scalable tools for investigating and understanding evaluation results.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗