CareerPlanGet AI match score →

Research Scientist - RL Training

Redwood City💼 Full-time🗓 2026-07-07 → 2026-07-31

Core

Foundational research role focused on generating data, reward signals, and training procedures to steer large language model (LLM) behavior using reinforcement learning.

Role type

Research Scientist (RL Training & Alignment)

Builds

Preference datasets, reward models, RL-ready corpora, and data pipelines for frontier AI labs.

Domain

Generative AI, Reinforcement Learning, Large Language Models

Deliverable

production ML models

Required skills

Reinforcement learning (RLHF, RLAIF, DPO, reward modeling), Python, PyTorch, HuggingFace, distributed training infrastructure, software engineering for research prototypes

Preferred skills

Training/fine-tuning 30B+ LLMs, Verl and SkyRL frameworks, ML infrastructure (AWS, GCP, Kubernetes, Slurm), Ph.D. in ML/RL

Technologies

PyTorch, HuggingFace, Verl, SkyRL, AWS, GCP, Kubernetes, Slurm

Responsibilities

Research and implement RL techniques to translate into data products; Design and build data pipelines for high-quality training signals; Prototype end-to-end RL training recipes; Collaborate with engineering and delivery teams to productize RL research; Stay current with large-scale LLM training and alignment research; Contribute to research publications.

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗