Senior AI Infrastructure Software Engineer - DGX Cloud
Core
Design, build, and maintain AI platforms enabling large-scale AI training, inferencing, fine-tuning, and Agentic AI in production.
Role type
Senior IC AI infrastructure software engineer
Builds
Scalable AI infrastructure services and platforms for NVIDIA DGX Cloud
Domain
Cloud infrastructure + AI/ML systems
Deliverable
production ML models | infrastructure
Required skills
Large-scale distributed systems, debugging/triage (app to hardware), Kubernetes, observability (ELK, Prometheus, Loki), Golang, Python, C/C++
Preferred skills
NVIDIA GPU/network tech (RDMA, IB, NCCL), DL frameworks (PyTorch, TensorFlow, JAX, Dynamo, Ray), datacenter-scale failure analysis
Technologies
Kubernetes, ELK, Prometheus, Loki, Golang, Python, C/C++, PyTorch, TensorFlow, JAX, Dynamo, Ray
Responsibilities
Develop platform/tools for large-scale AI/LLM/GenAI infrastructure; Optimize AI/ML workload efficiency and resiliency; Root cause analyze and triage failures from application to hardware level; Enhance infrastructure underpinning NVIDIA AI platforms; Co-design and implement APIs for resiliency stacks; Define reliability metrics to track and improve system/service reliability
Seniority
Senior, hands-on IC