CareerPlanGet AI match score →

Member of Technical Staff - Post-Training and RL

Palo Alto, CA💼 Full-time💰 $180,000–$180,000🗓 2026-05-26 → 2026-08-01

Core

Solving critical post-training and reinforcement learning challenges to improve AI models' reasoning, truthfulness, and real-world capabilities.

Role type

Post-training and RL engineer

Builds

AI models optimized via reward modeling, preference optimization (RLHF/DPO), and RL for capability improvement

Domain

Artificial Intelligence / Machine Learning

Deliverable

production ML models

Required skills

reinforcement learning, reward modeling, preference optimization (RLHF/DPO), model training, alignment methods

Preferred skills

experience with post-training, RLHF, or training models used by millions

Technologies

N/A

Responsibilities

Implement reward modeling, preference optimization (RLHF/DPO), and RL techniques to enhance model reasoning and truthfulness

Seniority

Individual Contributor

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗