CareerPlanGet AI match score →

Software Engineer, RL Training Infra

San Francisco💼 Full-time🗓 2026-05-23 → 2026-07-31

Core

Building and maintaining the infrastructure for large-scale reinforcement learning training runs for frontier AI models (e.g., o1, 5.5) used in products like Codex and ChatGPT.

Role type

Senior IC machine-learning infrastructure engineer (RL training)

Builds

Scalable, reliable RL training systems and orchestration layers for agentic models

Domain

Artificial Intelligence / Reinforcement Learning / Distributed Systems

Deliverable

production ML models

Required skills

Distributed systems debugging, RL training systems, scaling and orchestration, inference optimization, numerical problem solving, hardware failure diagnosis, system reliability engineering, performance optimization, multi-agent system integration, rapid prototyping

Preferred skills

Async RL systems, high-throughput ML infrastructure, GPU networking, production-critical infrastructure, research collaboration

Technologies

RL training frameworks, distributed computing stacks, GPU clusters, orchestration systems, serving systems, agent harnesses

Responsibilities

Debug urgent engineering and infrastructure issues across training, inference, and orchestration layers; Solve hard technical problems at the boundary of research and engineering; Improve reliability and efficiency of large-scale RL training runs; Turn recurring operational issues into robust tools and abstractions; Support researchers developing infra-heavy integrations like multi-agent capabilities; Collaborate with research and partner teams during tight model run timelines

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗