CareerPlanSign in

Member of Technical Staff - Evaluations

San Francisco, CA💼 Full-time🗓 2025-12-17 → 2026-09-25

Core

Conduct comparative analysis and build evaluation systems to measure LLM capabilities, reasoning, and alignment.

Role type

Senior IC machine-learning engineer (evaluation)

Builds

Generalizable evaluation frameworks and feedback loops for model improvement

Domain

Artificial Intelligence / Large Language Models

Deliverable

production ML models

Required skills

statistical analysis, experimental design, LLM evaluation methodologies, synthetic eval design, human feedback integration

Preferred skills

agentic task evaluation, real-world interaction data analysis

Technologies

LLMs, synthetic data, human feedback systems

Responsibilities

Conduct critical comparative analysis to advance understanding of model capabilities; Build and refine evaluation systems creating feedback loops between data, evals, and model behavior; Develop generalizable evaluation frameworks for reasoning, alignment, and usefulness; Collaborate with pre-training and post-training teams to translate insights into model improvements; Push boundaries of measurable metrics from synthetic evals to real-world interaction data

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.