CareerPlanGet AI match score →

Research Engineer, Infrastructure, RL Systems

San Francisco💼 Full-time💰 $350,000–$350,000🗓 2026-05-04 → 2026-07-31

Core

Design and build core infrastructure systems enabling scalable, efficient training of large models through reinforcement learning.

Role type

Senior IC infrastructure research engineer (RL systems)

Builds

Distributed RL training pipelines, rollout/reward systems, evaluation benchmarks, observability tools

Domain

AI/ML infrastructure, Reinforcement Learning, Large-scale distributed systems

Deliverable

production ML models

Required skills

Deep learning frameworks (PyTorch, JAX), distributed training, cluster orchestration, observability, code optimization, system reliability

Preferred skills

Large-scale LLM training (10B+ params), RL workloads (PPO, DPO, RLHF), high-performance engineering, open-source contributions

Technologies

Kubernetes, Slurm, Prometheus, Grafana, OpenTelemetry

Responsibilities

Design and optimize infrastructure for large-scale RL and post-training workloads; Improve reliability and scalability of RL training pipelines; Develop shared monitoring and observability tools; Collaborate with researchers to translate algorithmic ideas into production-grade pipelines; Build evaluation and benchmarking infrastructure; Publish learnings via documentation or open-source

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗