CareerPlanSign in

Senior Software Engineer - Managed Kubernetes

San Francisco Office (Fremont St)💼 Full-time🗓 2026-09-24 → 2026-09-25

Core

Design, build, and maintain scalable control plane services, operators, and custom Kubernetes controllers for a managed Kubernetes platform purpose-built for AI workloads running on bare metal.

Role type

Senior Software Engineer (Managed Kubernetes Infrastructure)

Builds

Managed Kubernetes platform, Managed Slurm on Kubernetes, and higher-level platform services for inference and AIOps

Domain

AI Cloud Infrastructure, Distributed Systems, GPU-accelerated Computing

Deliverable

production ML models | infrastructure

Required skills

Kubernetes internals (controllers, schedulers, operators, CRDs, CSI, CNI), Go, Python, Linux systems, networking, distributed systems fundamentals, observability (Prometheus, Grafana, distributed tracing)

Preferred skills

Managed Kubernetes services (GKE, EKS, AKS), NVIDIA GPU/networking ecosystem (GPU Operator, DCGM, MIG, NCCL), HPC and job schedulers (Slurm, KAI, Volcano, Kueue), GPU/InfiniBand/RDMA, storage architecture for AI/ML

Technologies

Go, Python, Kubernetes, Cilium, Multus, InfiniBand, RoCE, RDMA, Prometheus, Grafana, NVIDIA GPU stack

Responsibilities

Design and build scalable control plane services and custom Kubernetes controllers; develop automation for cluster lifecycle management; build GPU-aware orchestration systems; partner on networking solutions for AI workloads; write resilient systems for failure handling; develop platform services for inference and model serving; build internal tools and CLIs for ML/AI teams; support and debug production issues

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.