Senior Production Engineer - DGX Cloud
Core
Scaling AI infrastructure by building and managing large-scale GPU clusters for diverse AI workloads.
Role type
Senior Production Engineer (SRE/DevOps)
Builds
Custom software for GPU asset provisioning, configuration, and lifecycle management across cloud providers; monitoring and health management systems for GPU clusters.
Domain
AI computing, GPU deep learning, cloud infrastructure
Deliverable
production ML models | infrastructure
Required skills
Site reliability engineering, incident management, production system observability, automated deployments, systems programming (Go, Python), data structures and algorithms
Preferred skills
Managing and automating large-scale distributed systems independent of cloud providers, cluster management systems (Kubernetes, Slurm, Bright Cluster Manager), operational excellence in AI infrastructure
Technologies
Kubernetes, Slurm, Bright Cluster Manager, Go, Python
Responsibilities
Implement monitoring and health management capabilities for GPU assets; evaluate system failures and improve services based on incident management processes; work on custom software for GPU asset lifecycle management across cloud providers.
Seniority
Senior, hands-on IC