Capacity & Efficiency Infrastructure
Core
Design, implement, test, and optimize distributed training infrastructure for large-scale GPU clusters, focusing on efficiency, telemetry, and performance bottlenecks.
Role type
Senior IC infrastructure engineer (distributed training & HPC)
Builds
Distributed training infrastructure, telemetry systems, and optimization tools for ML fleets
Domain
High-performance computing, large-scale machine learning, generative AI
Deliverable
production ML models
Required skills
C++, Python, CUDA, NCCL, PyTorch, JAX, GPU architecture fundamentals, distributed computing profiling, InfiniBand/NVLink networking, low-level GPU programming, architectural leadership
Preferred skills
Experience with next-generation accelerators (NVIDIA, MAIA), collective communication library optimization
Technologies
Python, C++, CUDA, Triton, NCCL, PyTorch, JAX, InfiniBand, NVLink
Responsibilities
Design and optimize distributed training infrastructure; build telemetry systems for infrastructure and model performance; profile and debug bottlenecks across compute, memory, and networking; drive architectural improvements for ML services; build tools for fleet-wide efficiency insights; optimize collective communication libraries; partner with researchers and hardware teams
Seniority
Senior, hands-on IC