Staff Software Engineer - Managed Kubernetes
Core
Design and build a purpose-built Managed Kubernetes platform for AI workloads running on bare metal, enabling scalable training and inference for hyperscalers and enterprises.
Role type
Staff Software Engineer (Infrastructure/Orchestration)
Builds
Managed Kubernetes control plane, Managed Slurm on Kubernetes, GPU-aware orchestration services, and multi-tenant inference platforms.
Domain
AI Cloud Infrastructure, Distributed Systems, GPU Computing
Deliverable
production ML models | infrastructure
Required skills
Kubernetes internals (API, controllers, schedulers, operators, CRDs, CSI, CNI), Go, Python, GPU orchestration (NVIDIA GPU Operator, DCGM, MIG, time-slicing), distributed systems, Linux networking (L2-L7, RDMA, InfiniBand), observability, infrastructure-as-code
Preferred skills
Managed K8s services (GKE, EKS, AKS), NVIDIA Network Operator, NCCL tuning, HPC schedulers (Slurm, Kueue), confidential computing, CNCF contributions
Technologies
Kubernetes, Go, Python, NVIDIA GPU Operator, DCGM, NCCL, InfiniBand, RoCE, RDMA, Cilium, Multus, Prometheus, Grafana
Responsibilities
Drive technical vision for bare-metal Managed Kubernetes; Integrate NVIDIA open-source ecosystem for GPU workloads; Design GPU-aware orchestration and self-healing systems; Lead chaos engineering and operational excellence; Serve as technical bridge between Orchestration, Network, Storage, and Security teams; Mentor engineers and set technical direction.
Seniority
Staff, technical leadership & hands-on IC