Research Engineer / Performance Engineer, RL Distributed Systems
Core
Design, build, and operate distributed systems that run reinforcement learning at scale, ensuring correctness under failure and maximizing compute efficiency.
Role type
Senior IC Research Engineer (Distributed Systems)
Builds
Distributed infrastructure for RL training, sampling, and environment execution
Domain
AI/ML infrastructure, Distributed Systems
Deliverable
production ML models
Required skills
Python, Rust/C++/Go, distributed systems fundamentals, fault tolerance, autoscaling, observability, debugging complex failures
Preferred skills
ML training/inference infrastructure, Kubernetes, sandboxed execution, high-performance networking (RDMA), async Python (Trio/asyncio), RL workloads
Technologies
Python, Rust, C++, Go, Kubernetes, RDMA, Trio, asyncio
Responsibilities
Design schedulers for heterogeneous clusters, build failure detection and recovery systems, scale environment execution, design autoscaling policies, build diagnostics systems, trace data corruption bugs, design operational interfaces for automated tools
Seniority
Senior, hands-on IC
