Software Engineer, RL Training Infra
Core
Building and maintaining the infrastructure for large-scale reinforcement learning training runs for frontier AI models (e.g., o1, 5.5) used in products like Codex and ChatGPT.
Role type
Senior IC machine-learning infrastructure engineer (RL training)
Builds
Scalable, reliable RL training systems and orchestration layers for agentic models
Domain
Artificial Intelligence / Reinforcement Learning / Distributed Systems
Deliverable
production ML models
Required skills
Distributed systems debugging, RL training systems, scaling and orchestration, inference optimization, numerical problem solving, hardware failure diagnosis, system reliability engineering, performance optimization, multi-agent system integration, rapid prototyping
Preferred skills
Async RL systems, high-throughput ML infrastructure, GPU networking, production-critical infrastructure, research collaboration
Technologies
RL training frameworks, distributed computing stacks, GPU clusters, orchestration systems, serving systems, agent harnesses
Responsibilities
Debug urgent engineering and infrastructure issues across training, inference, and orchestration layers; Solve hard technical problems at the boundary of research and engineering; Improve reliability and efficiency of large-scale RL training runs; Turn recurring operational issues into robust tools and abstractions; Support researchers developing infra-heavy integrations like multi-agent capabilities; Collaborate with research and partner teams during tight model run timelines
Seniority
Senior, hands-on IC
