Research Scientist - RL Training
Core
Foundational research role focused on generating data, reward signals, and training procedures to steer large language model (LLM) behavior using reinforcement learning.
Role type
Research Scientist (RL Training & Alignment)
Builds
Preference datasets, reward models, RL-ready corpora, and data pipelines for frontier AI labs.
Domain
Generative AI, Reinforcement Learning, Large Language Models
Deliverable
production ML models
Required skills
Reinforcement learning (RLHF, RLAIF, DPO, reward modeling), Python, PyTorch, HuggingFace, distributed training infrastructure, software engineering for research prototypes
Preferred skills
Training/fine-tuning 30B+ LLMs, Verl and SkyRL frameworks, ML infrastructure (AWS, GCP, Kubernetes, Slurm), Ph.D. in ML/RL
Technologies
PyTorch, HuggingFace, Verl, SkyRL, AWS, GCP, Kubernetes, Slurm
Responsibilities
Research and implement RL techniques to translate into data products; Design and build data pipelines for high-quality training signals; Prototype end-to-end RL training recipes; Collaborate with engineering and delivery teams to productize RL research; Stay current with large-scale LLM training and alignment research; Contribute to research publications.
Seniority
Senior, hands-on IC