Agent Post-Training, Personality
Core
Define agent personality traits (thoughtfulness, clarity, proactivity) and translate them into evals, training data, and reward signals to improve OpenAI's frontier agents.
Role type
Senior IC machine-learning engineer (post-training personality)
Builds
Training data, graders, reward models, and RL objectives for agent collaboration behaviors
Domain
AI research and deployment (LLMs, post-training, RLHF)
Deliverable
production ML models
Required skills
Machine learning, statistics, behavioral science, HCI, LLMs, post-training, RL/RLHF, reward modeling, evals, synthetic data, production ML systems
Preferred skills
Product thinking, research, communication taste, hypothesis generation, data pipeline building
Technologies
LLMs, RLHF, reward models, synthetic data pipelines
Responsibilities
Develop hypotheses from qualitative judgments about model behavior, study user signals to understand trust and satisfaction, produce preference data with human experts, improve reward models and RL objectives, partner with product teams to validate improvements, own projects end-to-end from observation to launch
Seniority
Senior, hands-on IC