CareerPlanSign in

Pre-Training Infrastructure

United States, Multiple Locations, Multiple Locations💼 Full-time🗓 2025-11-25 → 2026-09-25

Core

Design, implement, test, and optimize distributed training infrastructure for large-scale GPU clusters to support pretraining compute roadmaps and generative AI workloads.

Role type

Senior IC distributed systems engineer (GPU infrastructure)

Builds

Distributed training infrastructure, collective communication libraries, and high-performance computing systems for large-scale AI models.

Domain

High-performance computing, distributed systems, GPU programming, generative AI infrastructure

Deliverable

production ML models | infrastructure

Required skills

C++, Python, CUDA, NCCL, distributed computing, large-scale systems, performance profiling, benchmarking, networking (InfiniBand, NVLink), storage systems, parallelism strategies

Preferred skills

Experience with NVIDIA/AMD accelerators, architectural decision-making, leading technical projects

Technologies

Python, C++, CUDA, NCCL, PyTorch, InfiniBand, NVLink

Responsibilities

Profile, benchmark, and debug performance bottlenecks across compute, memory, networking, and storage subsystems; Optimize collective communication libraries for emerging topologies; Collaborate with hardware teams to optimize for next-generation accelerators; Gather data and insights to develop the pretraining compute roadmap.

Seniority

Senior, hands-on IC

Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.