Principal Software Engineer, Distributed Systems Engineer - DGX Cloud
Core
Designing and operating large-scale GPU clusters and AI infrastructure on Kubernetes to support diverse AI workloads.
Role type
Principal Software Engineer, Distributed Systems
Builds
Production AI clusters and custom Kubernetes scheduling software for GPU resources
Domain
AI computing, GPU infrastructure, cloud-native systems
Deliverable
production ML models | infrastructure
Required skills
Kubernetes API development, cluster operations, GPU resource scheduling, systems programming (Go/Python), data structures and algorithms, incident management, large-scale distributed systems architecture
Preferred skills
Experience with Slurm or Bright Cluster Manager, automating distributed systems independent of cloud providers, operational excellence in AI infrastructure
Technologies
Kubernetes, Go, Python, Slurm, Bright Cluster Manager
Responsibilities
Develop custom software for GPU resource scheduling on Kubernetes, implement monitoring and health management for GPU assets, evaluate system failures and improve services via incident management processes
Seniority
Principal, hands-on IC with strategy & mentorship