Senior DGX Cloud AI Infrastructure Software Engineer
Core
Design, build, and maintain AI infrastructure enabling large-scale AI training and inferencing for researchers.
Role type
Senior IC AI infrastructure software engineer
Builds
Infrastructure software and tools for large-scale AI, LLM, and GenAI systems
Domain
High-performance computing and AI infrastructure
Deliverable
production ML models | infrastructure
Required skills
Large-scale distributed systems development, root cause analysis across application to hardware levels, observability platform operation, Python, C/C++, software design
Preferred skills
Large-scale AI cluster experience, NVIDIA GPU and network technologies (RDMA, IB, NCCL), deep learning frameworks (PyTorch, TensorFlow, JAX, Ray), datacenter-scale failure analysis
Technologies
Python, C/C++, ELK, Prometheus, Loki, PyTorch, TensorFlow, JAX, Ray, RDMA, IB, NCCL
Responsibilities
Develop infrastructure software and tools for large-scale AI, LLM, and GenAI; optimize tools for infrastructure efficiency and resiliency; root cause and analyze failures from application to hardware levels; enhance infrastructure underpinning NVIDIA's AI platforms; co-design and implement APIs for resiliency stacks; define reliability metrics to track and improve system reliability
Seniority
Senior, hands-on IC