CareerPlanSign in

Distinguished Engineer, Production Engineering, Data Center Automation

US, CA, Santa Clara💼 Full-time💰 $320,000–$320,000🗓 2026-09-02 → 2026-09-25

Core

Senior technical leader defining architectural direction and operational frameworks for large-scale DGX Cloud GPU infrastructure across on-prem, hyperscaler, and cloud partner environments.

Role type

Distinguished Engineer, Production Engineering (Cluster Operations & Strategy)

Builds

Operational frameworks, workflows, and interfaces for DGX Cloud capacity management and reliability.

Domain

Cloud Infrastructure / GPU Systems / Distributed Systems

Deliverable

production ML models | infrastructure

Required skills

Large-scale distributed systems architecture, Kubernetes service management, cross-organizational technical leadership, production operating model design, infrastructure automation, release and runtime preparation, service reliability coordination.

Preferred skills

Production operating model development for multi-platform environments, setting company-wide operating standards and APIs, building automation interfaces connecting platform and infrastructure teams.

Technologies

Kubernetes, DGX Cloud, Hyperscaler platforms, NeoCloud, Bare-metal infrastructure.

Responsibilities

Define long-range technical strategy for operating DGX Cloud clusters; Define architectural vision for cluster lifecycle and steady-state operability; Guide roadmap for cross-organizational investments in production readiness; Make high-impact technical decisions on platform, hardware, and service coordination; Develop robust workflows and engineering collaboration across infrastructure and service domains.

Seniority

Distinguished, hands-on IC with strategic leadership

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.