CareerPlanGet AI match score →

Researcher, Context - Agent Post-Training

San Francisco💼 Full-time💰 $250,000–$250,000🗓 2026-05-22 → 2026-07-31

Core

Scaling compute spent on context to enable the next paradigm of model training for frontier agents (Codex, ChatGPT) that operate computers and collaborate with people.

Role type

Senior IC machine-learning researcher (agent post-training)

Builds

Frontier training stack, RL pipelines, graders, reward signals, evals, diagnostics, and production agent harness.

Domain

AI research, large-scale model training, agent systems, computer use

Deliverable

production ML models

Required skills

machine learning fundamentals, software engineering, statistics, LLMs, RL, RLHF/RLAIF, post-training, evals, graders, synthetic data, model training, coding agents, tool-using agents, production ML systems

Preferred skills

research taste, engineering execution, product impact focus, cross-functional collaboration, building load-bearing systems

Technologies

RL, RLHF, RLAIF, synthetic data, production ML systems

Responsibilities

Design and run experiments to improve scaling of compute on context; Own end-to-end improvements to the post-training stack including RL, data pipelines, graders, reward signals, evals, diagnostics, and model-behavior analysis; Build evals and environments that expose model failures and turn them into training data or product fixes; Partner with product teams to translate user needs into model improvements; Work on early-training and alignment interventions including data mixtures, objectives, synthetic data, and eval loops; Decide which integrations and fixes are ready for major model runs; Improve machinery for large-scale training including experiment velocity, reliability, observability, reproducibility, cost, latency, and production readiness; Debug hard failures in shipped models and turn qualitative behavior into concrete hypotheses and fixes

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗