AI ⽹络集合通信库运营⼯程师
Core
Operate and optimize collective communication libraries (e.g., NCCL) for large-scale AI training and inference clusters, ensuring high availability and performance.
Role type
Senior IC AI infrastructure engineer (collective communication)
Builds
High-performance communication services for distributed AI workloads
Domain
AI Infrastructure / High-Performance Computing
Deliverable
production ML models
Required skills
Collective communication algorithms (AllReduce, AllGather, Ring, Tree), NCCL architecture and tuning, GPU interconnect mechanisms (NVLink, RDMA, InfiniBand, RoCE), Linux systems administration, Python/Go/Bash programming, AI parallelism strategies (TP, PP, DP, EP, CP)
Preferred skills
CUDA programming, DeepEP/Gloo/MSCCL experience, automated observability tool development
Technologies
NCCL, NVLink, InfiniBand, RoCE, DCGM, Nsight Systems, nccl-tests
Responsibilities
Deploy, configure, and upgrade collective communication libraries; monitor and optimize communication bandwidth and latency; diagnose and resolve training hangs and communication timeouts; collaborate with framework teams on parallel strategy tuning; develop diagnostic and automation tools; plan capacity and evaluate communication library versions for new hardware.
Seniority
Senior, hands-on IC
