Senior Software Engineer, DGX Cloud Production Engineering
Core
Build and operate automation, tooling, and operational systems for large-scale GPU clusters to ensure reliability, scalability, and safety for AI research and production workloads.
Role type
Senior IC infrastructure engineer (GPU clusters, Kubernetes)
Builds
Automation, provisioning, validation, monitoring, and lifecycle management tools for GPU clusters
Domain
Cloud infrastructure + High-Performance Computing (GPU)
Deliverable
production ML models | infrastructure
Required skills
Python, Go, Linux, Kubernetes, containers, cloud infrastructure, infrastructure automation, distributed systems troubleshooting
Preferred skills
GPU infrastructure, Kubernetes operators, GitOps, Terraform, ArgoCD, fleet automation, SLOs, incident response, observability, reliability practices, BMaaS, VMaaS, managed Kubernetes, multi-cloud infrastructure
Responsibilities
Build automation for large-scale GPU clusters across cloud and on-prem environments; Develop tools for cluster provisioning, validation, upgrades, monitoring, and repair; Improve cluster bringup and production workflows; Reduce manual touches via APIs and GitOps; Participate in on-call and incident response; Partner with platform, storage, networking, and security teams
Seniority
Senior, hands-on IC