Staff Software Engineer, RL Environments
Core
Design and build the technical foundation for creating, running, verifying, and delivering reinforcement learning (RL) environments at scale, including sandboxed execution, rollout orchestration, and verifier frameworks.
Role type
Staff Software Engineer (RL Environments)
Builds
Scalable RL environment platforms, task suites, graders, and authoring surfaces for engineers and domain experts.
Domain
Artificial Intelligence / Reinforcement Learning / Distributed Systems
Deliverable
production ML models
Required skills
Distributed systems design, Python, containerization (Docker, VMs, gVisor/Firecracker), Kubernetes, high-throughput backend systems, LLM agent loops, tool calling, reward modeling
Preferred skills
RLHF/RLAIF/RLVR algorithms, reward hacking defense, RL training stacks (verl, TRL, Ray, vLLM), cloud-native infrastructure (AWS/GCP/Azure), observability, internal tooling for non-engineers
Technologies
Python, TypeScript, React, Go, Rust, Docker, Kubernetes, AWS, GCP, Azure, Ray, vLLM, SGLang
Responsibilities
Design platform architecture for sandboxed execution and environment versioning; build graders that withstand adversarial optimization; instrument real applications to design task suites exposing capability gaps; set technical direction across teams while writing core code.
Seniority
Staff, hands-on IC with technical leadership