CareerPlanGet AI match score →

Senior Software Engineer, DGX Cloud AI Infrastructure

5 Locations💼 Full-time💰 $184,000–$184,000🗓 2026-06-04 → 2026-07-31

Core

Lead the bring-up, triage, benchmarking, analysis, and optimization of distributed training and inference workloads for large-scale LLMs across NVIDIA GPU platforms.

Role type

Senior IC distributed systems engineer (AI infrastructure)

Builds

High-performance distributed training and inference stacks for large language models

Domain

Generative AI, HPC, GPU clusters, distributed computing

Deliverable

production ML models

Required skills

Expert-level Python and C/C++, deep debugging of multi-GPU/multi-node workloads, NCCL and CUDA-aware distributed execution, profiling and optimization of compute/memory/networking layers, scaling efficiency analysis (data/tensor/pipeline/expert parallelism), root-cause analysis of cluster failures, building resilience and failure-attribution systems, technical leadership and mentorship

Preferred skills

RDMA software stack (IB verbs, UCX, libfabric), GPU cluster fabrics (NVLink, NVSwitch, PCIe, RoCE, InfiniBand), building benchmark harnesses and qualification tooling

Technologies

PyTorch, NeMo, Megatron, TensorRT-LLM, Nsight Systems, NCCL tests

Responsibilities

Lead bring-up and validation of large-scale AI clusters; tune and benchmark pre-training/post-training/inference workloads; profile and optimize end-to-end workload performance; analyze scaling efficiency for distributed LLMs; own root-cause analysis of complex failures; build resilience and failure-attribution stacks; define and build repeatable benchmark suites; tune runtime settings and deployment configurations; mentor engineers and drive technical standards

Seniority

Senior, hands-on IC with technical leadership

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗