Agent Post-Training Research
Core
Improving the capabilities, reliability, and product fit of OpenAI's agentic models by designing experiments, building training infrastructure, and creating evaluation systems for models that act in the world.
Role type
Senior IC machine-learning researcher (agentic systems)
Builds
Frontier AI agents capable of coding, tool use, computer operation, and multi-agent collaboration
Domain
Artificial Intelligence / Machine Learning / Agentic Systems
Deliverable
production ML models
Required skills
machine learning fundamentals, software engineering, statistics, LLMs, reinforcement learning, RLHF/RLAIF, post-training, evaluation design, synthetic data generation, model debugging, cross-functional collaboration
Preferred skills
experience with coding agents, tool-using agents, production ML systems, open-ended problem solving, product impact focus
Technologies
RL, RLHF, RLAIF, synthetic data pipelines, evaluation frameworks, large-scale training infrastructure
Responsibilities
Design and run experiments to improve agentic model behavior across coding, tool use, and multi-agent collaboration; Own end-to-end improvements to the post-training stack including RL, data pipelines, and graders; Build evals and environments to expose model failures and convert them into training data; Partner with product teams to translate user needs into model improvements; Work on early-training and alignment interventions; Improve machinery for large-scale training and launch; Debug hard failures in shipped models
Seniority
Senior, hands-on IC