CareerPlanSign in

Research Scientist (Remote/US/LATAM)

Fully Remote🌐 Remote💼 Full-time🗓 2026-07-06 → 2026-09-26

Core

Design and build frontier-grade evaluation benchmarks and packages for LLMs, focusing on reasoning, coding, agents, and tool use to measure model capabilities.

Role type

Research Scientist (LLM Evaluations & Benchmarking)

Builds

Expert-verified benchmark datasets, evaluation packages, and public research papers.

Domain

Artificial Intelligence / Large Language Models / Evaluation Metrics

Deliverable

production ML models

Required skills

LLM benchmarking, evaluation research, construct validity, psychometrics, rubric design, expert recruitment, research writing, coding evaluation, agentic evaluation

Preferred skills

Spanish fluency, safety evaluation expertise

Technologies

LLMs, multi-modal models, deterministic verifiers

Responsibilities

Design original benchmark structures, build evaluation packages with ground truth, recruit and calibrate expert pools, act as technical liaison for research labs, deliver sample packages and publish research papers

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.