CareerPlanGet AI match score →

Principal Software Engineer, Distributed Systems Engineer - DGX Cloud

2 Locations💼 Full-time💰 $272,000–$272,000🗓 2026-06-25 → 2026-07-31

Core

Designing and operating large-scale GPU clusters and AI infrastructure on Kubernetes to support diverse AI workloads.

Role type

Principal Software Engineer, Distributed Systems

Builds

Production AI clusters and custom Kubernetes scheduling software for GPU resources

Domain

AI computing, GPU infrastructure, cloud-native systems

Deliverable

production ML models | infrastructure

Required skills

Kubernetes API development, cluster operations, GPU resource scheduling, systems programming (Go/Python), data structures and algorithms, incident management, large-scale distributed systems architecture

Preferred skills

Experience with Slurm or Bright Cluster Manager, automating distributed systems independent of cloud providers, operational excellence in AI infrastructure

Technologies

Kubernetes, Go, Python, Slurm, Bright Cluster Manager

Responsibilities

Develop custom software for GPU resource scheduling on Kubernetes, implement monitoring and health management for GPU assets, evaluate system failures and improve services via incident management processes

Seniority

Principal, hands-on IC with strategy & mentorship

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗