Member of Technical Staff - Research Software Engineer
Core
Bridge the gap between research and production by designing and optimizing scalable training infrastructure for frontier AI models, including reinforcement learning loops and massive-scale data pipelines.
Role type
Senior IC research software engineer (distributed training infrastructure)
Builds
Scalable training systems for reinforcement learning and distributed GPU training
Domain
Artificial Intelligence / Distributed Systems / High-Performance Computing
Deliverable
production ML models
Required skills
Distributed training & inference, Data infrastructure, PyTorch, JAX, GPU parallelism, Communication optimization, Containerization, Experiment orchestration
Preferred skills
Triton/custom kernels, FSDP/ZeRO, NCCL/RDMA, Large-scale dataset curation
Technologies
PyTorch, JAX, Ray, Kubernetes, Slurm, Triton, NCCL, RDMA
Responsibilities
Design and optimize large-scale training loops and data pipelines; Implement state-of-the-art techniques ensuring numerical stability and efficiency; Build internal tooling for launching, monitoring, and reproducing experiments; Diagnose deep bottlenecks across the training stack; Translate research prototypes into production-grade infrastructure
Seniority
Senior, hands-on IC