Researcher, Context - Agent Post-Training
Core
Scaling compute spent on context to enable the next paradigm of model training for frontier agents (Codex, ChatGPT) that operate computers and collaborate with people.
Role type
Senior IC machine-learning researcher (agent post-training)
Builds
Frontier training stack, RL pipelines, graders, reward signals, evals, diagnostics, and production agent harness.
Domain
AI research, large-scale model training, agent systems, computer use
Deliverable
production ML models
Required skills
machine learning fundamentals, software engineering, statistics, LLMs, RL, RLHF/RLAIF, post-training, evals, graders, synthetic data, model training, coding agents, tool-using agents, production ML systems
Preferred skills
research taste, engineering execution, product impact focus, cross-functional collaboration, building load-bearing systems
Technologies
RL, RLHF, RLAIF, synthetic data, production ML systems
Responsibilities
Design and run experiments to improve scaling of compute on context; Own end-to-end improvements to the post-training stack including RL, data pipelines, graders, reward signals, evals, diagnostics, and model-behavior analysis; Build evals and environments that expose model failures and turn them into training data or product fixes; Partner with product teams to translate user needs into model improvements; Work on early-training and alignment interventions including data mixtures, objectives, synthetic data, and eval loops; Decide which integrations and fixes are ready for major model runs; Improve machinery for large-scale training including experiment velocity, reliability, observability, reproducibility, cost, latency, and production readiness; Debug hard failures in shipped models and turn qualitative behavior into concrete hypotheses and fixes
Seniority
Senior, hands-on IC