Research Scientist / Engineer – Reinforcement Learning Infrastructure
Core
Design, build, and scale distributed reinforcement learning post-training systems that couple policy optimization with large fleets of inference workers, agentic environments, and reward/verification systems for frontier-scale LLMs.
Role type
Senior IC reinforcement learning infrastructure engineer
Builds
Distributed RL post-training systems, high-throughput rollout generation, RL environments for agentic tasks, and reward infrastructure (verifiers, LLM-as-judge pipelines)
Domain
Artificial Intelligence / Reinforcement Learning / Large Language Models
Deliverable
production ML models
Required skills
RL post-training (PPO/GRPO/RLHF/RLVR), distributed PyTorch training (FSDP, Tensor/Pipeline/Expert Parallel), RL environment design, reward function engineering, GPU cluster management, NCCL/MPI networking, vLLM/SGLang integration, Ray orchestration
Preferred skills
Asynchronous/disaggregated trainer-rollout architectures, Kubernetes orchestration for large fleets, research contributions to RL frameworks
Technologies
PyTorch, vLLM, SGLang, Ray, Kubernetes, NCCL, MPI, veRL, OpenRLHF, TRL
Responsibilities
Design distributed RL post-training systems orchestrating trainer, rollout, environment, and reward workloads; Build high-throughput rollout generation with inference engines and weight synchronization; Develop RL environments for agentic, multi-step tasks including sandboxed code execution; Build reward infrastructure including verifiable rewards and defenses against reward hacking; Develop evaluation, monitoring, and debugging tooling for large RL runs; Advance training efficiency and stability for production runs
Seniority
Senior, hands-on IC