Senior Manager, Kubernetes Runtime Engineering
Core
Lead the Runtime Engineering team to manage the full configuration lifecycle of NVIDIA Kubernetes Engine (NKE) tenant workload clusters, ensuring reliable and secure GPU workloads at scale.
Role type
Senior Manager, Kubernetes Runtime Engineering
Builds
Production-grade, multi-tenant Kubernetes platform with NVIDIA AI Container Runtime (AICR), GPU management operators, and cluster networking/storage components.
Domain
Cloud Infrastructure, Kubernetes, GPU Computing, AI Infrastructure
Deliverable
production ML models
Required skills
Kubernetes internals, cluster lifecycle management, security and compliance, API design, people management, cross-organization leadership
Preferred skills
NVIDIA GPU Operator, DCGM Exporter, NVLink-aware scheduling, hyperscale Kubernetes experience, open source contributions
Technologies
Kubernetes, AICR, CNI, CSI, Cluster API, kubeadm, RBAC, MIG, MPS
Responsibilities
Manage a team of engineers coordinating the container runtime stack; Drive architecture decisions for cluster networking, storage, and GPU resource partitioning; Define and implement cluster hardening standards and multi-tenancy isolation; Build and maintain tooling for AICR lifecycle management; Represent the runtime team in architecture reviews and customer communications.
Seniority
Senior, hands-on IC with people management