Member of Technical Staff - Post-Training and RL
Core
Solving critical post-training and reinforcement learning challenges to improve AI models' reasoning, truthfulness, and real-world capabilities.
Role type
Post-training and RL engineer
Builds
AI models optimized via reward modeling, preference optimization (RLHF/DPO), and RL for capability improvement
Domain
Artificial Intelligence / Machine Learning
Deliverable
production ML models
Required skills
reinforcement learning, reward modeling, preference optimization (RLHF/DPO), model training, alignment methods
Preferred skills
experience with post-training, RLHF, or training models used by millions
Technologies
N/A
Responsibilities
Implement reward modeling, preference optimization (RLHF/DPO), and RL techniques to enhance model reasoning and truthfulness
Seniority
Individual Contributor
Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.