CareerPlanSign in

Senior Software Engineer, DGX Cloud Production Engineering

US, CA, Santa Clara💼 Full-time💰 $184,000–$184,000🗓 2026-09-23 → 2026-09-25

Core

Build and operate automation, tooling, and operational systems for large-scale GPU clusters to ensure reliability, scalability, and safety for AI research and production workloads.

Role type

Senior IC infrastructure engineer (GPU clusters, Kubernetes)

Builds

Automation, provisioning, validation, monitoring, and lifecycle management tools for GPU clusters

Domain

Cloud infrastructure + High-Performance Computing (GPU)

Deliverable

production ML models | infrastructure

Required skills

Python, Go, Linux, Kubernetes, containers, cloud infrastructure, infrastructure automation, distributed systems troubleshooting

Preferred skills

GPU infrastructure, Kubernetes operators, GitOps, Terraform, ArgoCD, fleet automation, SLOs, incident response, observability, reliability practices, BMaaS, VMaaS, managed Kubernetes, multi-cloud infrastructure

Responsibilities

Build automation for large-scale GPU clusters across cloud and on-prem environments; Develop tools for cluster provisioning, validation, upgrades, monitoring, and repair; Improve cluster bringup and production workflows; Reduce manual touches via APIs and GitOps; Participate in on-call and incident response; Partner with platform, storage, networking, and security teams

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.