CareerPlanSign in

Member of Technical Staff - Distributed Training Engineer

San Francisco💼 Full-time🗓 2025-07-29 → 2026-09-25

Core

Design, implement, and optimize distributed training infrastructure for large-scale GPU clusters to power next-generation foundation models.

Role type

Senior IC distributed training systems engineer

Builds

Scalable distributed training infrastructure, data loading systems, and checkpointing mechanisms for GPU clusters

Domain

Artificial Intelligence / Distributed Systems / High-Performance Computing

Deliverable

production ML models

Required skills

Distributed training infrastructure (PyTorch DDP/FSDP, DeepSpeed ZeRO, Megatron-LM), Performance profiling and debugging, Hardware accelerator and networking topology knowledge, Data pipeline optimization for ML workloads

Preferred skills

Mixture of Experts (MoE) training, Large-scale distributed training (100+ GPUs), Open-source contributions to training infrastructure

Technologies

PyTorch, DeepSpeed, Megatron-LM, NCCL, GPU clusters

Responsibilities

Design and build core systems for fast and reliable large training runs; Build scalable distributed training infrastructure for GPU clusters; Implement and tune parallelism/sharding strategies; Optimize distributed efficiency (topology-aware collectives, comm/compute overlap, straggler mitigation); Build data loading systems to eliminate I/O bottlenecks; Develop checkpointing mechanisms balancing memory and recovery; Create monitoring, profiling, and debugging tools for training stability

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.