CareerPlanSign in

Research Scientist / Engineer – Reinforcement Learning Infrastructure

Redwood City, CA💼 Full-time🗓 2026-07-24 → 2026-09-26

Core

Design, build, and scale distributed reinforcement learning post-training systems that couple policy optimization with large fleets of inference workers, agentic environments, and reward/verification systems for frontier-scale LLMs.

Role type

Senior IC reinforcement learning infrastructure engineer

Builds

Distributed RL post-training systems, high-throughput rollout generation, RL environments for agentic tasks, and reward infrastructure (verifiers, LLM-as-judge pipelines)

Domain

Artificial Intelligence / Reinforcement Learning / Large Language Models

Deliverable

production ML models

Required skills

RL post-training (PPO/GRPO/RLHF/RLVR), distributed PyTorch training (FSDP, Tensor/Pipeline/Expert Parallel), RL environment design, reward function engineering, GPU cluster management, NCCL/MPI networking, vLLM/SGLang integration, Ray orchestration

Preferred skills

Asynchronous/disaggregated trainer-rollout architectures, Kubernetes orchestration for large fleets, research contributions to RL frameworks

Technologies

PyTorch, vLLM, SGLang, Ray, Kubernetes, NCCL, MPI, veRL, OpenRLHF, TRL

Responsibilities

Design distributed RL post-training systems orchestrating trainer, rollout, environment, and reward workloads; Build high-throughput rollout generation with inference engines and weight synchronization; Develop RL environments for agentic, multi-step tasks including sandboxed code execution; Build reward infrastructure including verifiable rewards and defenses against reward hacking; Develop evaluation, monitoring, and debugging tooling for large RL runs; Advance training efficiency and stability for production runs

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.