CareerPlanSign in

Staff Software Engineer - Managed Kubernetes

San Francisco Office (Fremont St)💼 Full-time🗓 2026-09-24 → 2026-09-25

Core

Design and build a purpose-built Managed Kubernetes platform for AI workloads running on bare metal, enabling scalable training and inference for hyperscalers and enterprises.

Role type

Staff Software Engineer (Infrastructure/Orchestration)

Builds

Managed Kubernetes control plane, Managed Slurm on Kubernetes, GPU-aware orchestration services, and multi-tenant inference platforms.

Domain

AI Cloud Infrastructure, Distributed Systems, GPU Computing

Deliverable

production ML models | infrastructure

Required skills

Kubernetes internals (API, controllers, schedulers, operators, CRDs, CSI, CNI), Go, Python, GPU orchestration (NVIDIA GPU Operator, DCGM, MIG, time-slicing), distributed systems, Linux networking (L2-L7, RDMA, InfiniBand), observability, infrastructure-as-code

Preferred skills

Managed K8s services (GKE, EKS, AKS), NVIDIA Network Operator, NCCL tuning, HPC schedulers (Slurm, Kueue), confidential computing, CNCF contributions

Technologies

Kubernetes, Go, Python, NVIDIA GPU Operator, DCGM, NCCL, InfiniBand, RoCE, RDMA, Cilium, Multus, Prometheus, Grafana

Responsibilities

Drive technical vision for bare-metal Managed Kubernetes; Integrate NVIDIA open-source ecosystem for GPU workloads; Design GPU-aware orchestration and self-healing systems; Lead chaos engineering and operational excellence; Serve as technical bridge between Orchestration, Network, Storage, and Security teams; Mentor engineers and set technical direction.

Seniority

Staff, technical leadership & hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.