Senior Performance Engineer - DGX Cloud
Core
Characterize, diagnose, and optimize end-to-end performance of large-scale AI workloads across compute, network, storage, and software stacks for NVIDIA's DGX Cloud.
Role type
Senior Performance Engineer (Infrastructure)
Builds
Scalable DGX Cloud systems and optimized AI workload performance
Domain
High-Performance Computing (HPC) / AI Infrastructure / GPU Systems
Deliverable
production ML models | infrastructure
Required skills
C++, Python, operating systems, computer architecture, distributed systems, performance engineering, benchmarking, profiling, optimization
Preferred skills
CUDA, GPU computing systems, GPU performance analysis, deep learning frameworks (PyTorch, JAX/XLA), large-scale AI clusters, distributed training and inference
Technologies
CUDA, PyTorch, JAX/XLA
Responsibilities
Analyze end-to-end performance of large-scale AI workloads; Design and execute rigorous performance studies to establish baselines and diagnose bottlenecks; Define performance evaluation methodologies and success metrics; Use profiling and observability to create actionable optimization plans; Partner with deep learning engineers and GPU architects to deliver improvements.
Seniority
Senior, hands-on IC