CareerPlanGet AI match score →

Senior Production Engineer - DGX Cloud

6 Locations💼 Full-time💰 $168,000–$168,000🗓 2026-07-06 → 2026-07-31

Core

Scaling AI infrastructure by building and managing large-scale GPU clusters for diverse AI workloads.

Role type

Senior Production Engineer (SRE/DevOps)

Builds

Custom software for GPU asset provisioning, configuration, and lifecycle management across cloud providers; monitoring and health management systems for GPU clusters.

Domain

AI computing, GPU deep learning, cloud infrastructure

Deliverable

production ML models | infrastructure

Required skills

Site reliability engineering, incident management, production system observability, automated deployments, systems programming (Go, Python), data structures and algorithms

Preferred skills

Managing and automating large-scale distributed systems independent of cloud providers, cluster management systems (Kubernetes, Slurm, Bright Cluster Manager), operational excellence in AI infrastructure

Technologies

Kubernetes, Slurm, Bright Cluster Manager, Go, Python

Responsibilities

Implement monitoring and health management capabilities for GPU assets; evaluate system failures and improve services based on incident management processes; work on custom software for GPU asset lifecycle management across cloud providers.

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗