CareerPlanSign in

Capacity & Efficiency Infrastructure

United States, Multiple Locations, Multiple Locations💼 Full-time🗓 2026-03-20 → 2026-09-25

Core

Design, implement, test, and optimize distributed training infrastructure for large-scale GPU clusters, focusing on efficiency, telemetry, and performance bottlenecks.

Role type

Senior IC infrastructure engineer (distributed training & HPC)

Builds

Distributed training infrastructure, telemetry systems, and optimization tools for ML fleets

Domain

High-performance computing, large-scale machine learning, generative AI

Deliverable

production ML models

Required skills

C++, Python, CUDA, NCCL, PyTorch, JAX, GPU architecture fundamentals, distributed computing profiling, InfiniBand/NVLink networking, low-level GPU programming, architectural leadership

Preferred skills

Experience with next-generation accelerators (NVIDIA, MAIA), collective communication library optimization

Technologies

Python, C++, CUDA, Triton, NCCL, PyTorch, JAX, InfiniBand, NVLink

Responsibilities

Design and optimize distributed training infrastructure; build telemetry systems for infrastructure and model performance; profile and debug bottlenecks across compute, memory, and networking; drive architectural improvements for ML services; build tools for fleet-wide efficiency insights; optimize collective communication libraries; partner with researchers and hardware teams

Seniority

Senior, hands-on IC

Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.