Member of Technical Staff - Distributed Training Engineer
Core
Design, implement, and optimize distributed training infrastructure for large-scale GPU clusters to power next-generation foundation models.
Role type
Senior IC distributed training systems engineer
Builds
Scalable distributed training infrastructure, data loading systems, and checkpointing mechanisms for GPU clusters
Domain
Artificial Intelligence / Distributed Systems / High-Performance Computing
Deliverable
production ML models
Required skills
Distributed training infrastructure (PyTorch DDP/FSDP, DeepSpeed ZeRO, Megatron-LM), Performance profiling and debugging, Hardware accelerator and networking topology knowledge, Data pipeline optimization for ML workloads
Preferred skills
Mixture of Experts (MoE) training, Large-scale distributed training (100+ GPUs), Open-source contributions to training infrastructure
Technologies
PyTorch, DeepSpeed, Megatron-LM, NCCL, GPU clusters
Responsibilities
Design and build core systems for fast and reliable large training runs; Build scalable distributed training infrastructure for GPU clusters; Implement and tune parallelism/sharding strategies; Optimize distributed efficiency (topology-aware collectives, comm/compute overlap, straggler mitigation); Build data loading systems to eliminate I/O bottlenecks; Develop checkpointing mechanisms balancing memory and recovery; Create monitoring, profiling, and debugging tools for training stability
Seniority
Senior, hands-on IC