Distinguished Engineer, Production Engineering, Data Center Automation
Core
Senior technical leader defining architectural direction and operational frameworks for large-scale DGX Cloud GPU infrastructure across on-prem, hyperscaler, and cloud partner environments.
Role type
Distinguished Engineer, Production Engineering (Cluster Operations & Strategy)
Builds
Operational frameworks, workflows, and interfaces for DGX Cloud capacity management and reliability.
Domain
Cloud Infrastructure / GPU Systems / Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Large-scale distributed systems architecture, Kubernetes service management, cross-organizational technical leadership, production operating model design, infrastructure automation, release and runtime preparation, service reliability coordination.
Preferred skills
Production operating model development for multi-platform environments, setting company-wide operating standards and APIs, building automation interfaces connecting platform and infrastructure teams.
Technologies
Kubernetes, DGX Cloud, Hyperscaler platforms, NeoCloud, Bare-metal infrastructure.
Responsibilities
Define long-range technical strategy for operating DGX Cloud clusters; Define architectural vision for cluster lifecycle and steady-state operability; Guide roadmap for cross-organizational investments in production readiness; Make high-impact technical decisions on platform, hardware, and service coordination; Develop robust workflows and engineering collaboration across infrastructure and service domains.
Seniority
Distinguished, hands-on IC with strategic leadership