Research Engineer, Infrastructure, RL Systems
Core
Design and build core infrastructure systems enabling scalable, efficient training of large models through reinforcement learning.
Role type
Senior IC infrastructure research engineer (RL systems)
Builds
Distributed RL training pipelines, rollout/reward systems, evaluation benchmarks, observability tools
Domain
AI/ML infrastructure, Reinforcement Learning, Large-scale distributed systems
Deliverable
production ML models
Required skills
Deep learning frameworks (PyTorch, JAX), distributed training, cluster orchestration, observability, code optimization, system reliability
Preferred skills
Large-scale LLM training (10B+ params), RL workloads (PPO, DPO, RLHF), high-performance engineering, open-source contributions
Technologies
Kubernetes, Slurm, Prometheus, Grafana, OpenTelemetry
Responsibilities
Design and optimize infrastructure for large-scale RL and post-training workloads; Improve reliability and scalability of RL training pipelines; Develop shared monitoring and observability tools; Collaborate with researchers to translate algorithmic ideas into production-grade pipelines; Build evaluation and benchmarking infrastructure; Publish learnings via documentation or open-source
Seniority
Senior, hands-on IC