Senior Software Engineer, DGX Cloud AI Infrastructure
Core
Lead the bring-up, triage, benchmarking, analysis, and optimization of distributed training and inference workloads for large-scale LLMs across NVIDIA GPU platforms.
Role type
Senior IC distributed systems engineer (AI infrastructure)
Builds
High-performance distributed training and inference stacks for large language models
Domain
Generative AI, HPC, GPU clusters, distributed computing
Deliverable
production ML models
Required skills
Expert-level Python and C/C++, deep debugging of multi-GPU/multi-node workloads, NCCL and CUDA-aware distributed execution, profiling and optimization of compute/memory/networking layers, scaling efficiency analysis (data/tensor/pipeline/expert parallelism), root-cause analysis of cluster failures, building resilience and failure-attribution systems, technical leadership and mentorship
Preferred skills
RDMA software stack (IB verbs, UCX, libfabric), GPU cluster fabrics (NVLink, NVSwitch, PCIe, RoCE, InfiniBand), building benchmark harnesses and qualification tooling
Technologies
PyTorch, NeMo, Megatron, TensorRT-LLM, Nsight Systems, NCCL tests
Responsibilities
Lead bring-up and validation of large-scale AI clusters; tune and benchmark pre-training/post-training/inference workloads; profile and optimize end-to-end workload performance; analyze scaling efficiency for distributed LLMs; own root-cause analysis of complex failures; build resilience and failure-attribution stacks; define and build repeatable benchmark suites; tune runtime settings and deployment configurations; mentor engineers and drive technical standards
Seniority
Senior, hands-on IC with technical leadership