Principal Software Engineer, Distributed Systems Engineer - DGX Cloud
Core
Designing and scaling production AI infrastructure for large GPU clusters using Kubernetes to support diverse AI workloads.
Role type
Principal Software Engineer, Distributed Systems
Builds
Scalable GPU clusters and custom Kubernetes scheduling software for AI applications
Domain
AI computing, GPU infrastructure, Cloud-native systems
Deliverable
production ML models | infrastructure
Required skills
Kubernetes API development, Cluster operations, GPU resource scheduling, Systems programming (Go/Python), Data structures and algorithms, Incident management, Large-scale distributed systems design
Preferred skills
Cluster management systems (Slurm, Bright Cluster Manager), Independent cloud-agnostic system management, Operational excellence in AI infrastructure
Responsibilities
Develop custom software for GPU resource scheduling on Kubernetes, Implement monitoring and health management for GPU assets, Evaluate system failures and improve services via incident management processes
Seniority
Principal, hands-on IC