CareerPlanSign in

Research Scientist - Frontier Benchmarks

San Francisco, CA (Hybrid)💼 Full-time🗓 2026-05-30 → 2026-09-26

Core

Design state-of-the-art datasets and benchmarks for frontier LLM training and evaluation, translating research insights into narratives for customers and go-to-market teams.

Role type

Research Scientist (Data-Centric AI / LLM Evaluation)

Builds

Frontier benchmarks, expert-curated datasets, and technical reports for customer-facing presentations.

Domain

Artificial Intelligence / Large Language Models (LLM) / Data Operations

Deliverable

production ML models | product features | dashboards & analysis | research | client delivery

Required skills

rigorous experimental design, LLM evaluation, NLP, data operations, cross-functional collaboration, technical storytelling

Preferred skills

Ph.D. in machine learning or NLP, track record in measuring data impact on model behavior, GTM strategy understanding

Technologies

LLMs, NLP frameworks, benchmarking tools

Responsibilities

Design datasets driving frontier model training and evaluation; Translate benchmark insights into compelling narratives for customers; Partner with product and engineering to inform company roadmap; Represent research externally via publications and conference talks

Seniority

Senior, hands-on IC with strategic influence

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.