Senior Software Engineer - Managed Kubernetes
Core
Design, build, and maintain scalable control plane services, operators, and custom Kubernetes controllers for a managed Kubernetes platform purpose-built for AI workloads running on bare metal.
Role type
Senior Software Engineer (Managed Kubernetes Infrastructure)
Builds
Managed Kubernetes platform, Managed Slurm on Kubernetes, and higher-level platform services for inference and AIOps
Domain
AI Cloud Infrastructure, Distributed Systems, GPU-accelerated Computing
Deliverable
production ML models | infrastructure
Required skills
Kubernetes internals (controllers, schedulers, operators, CRDs, CSI, CNI), Go, Python, Linux systems, networking, distributed systems fundamentals, observability (Prometheus, Grafana, distributed tracing)
Preferred skills
Managed Kubernetes services (GKE, EKS, AKS), NVIDIA GPU/networking ecosystem (GPU Operator, DCGM, MIG, NCCL), HPC and job schedulers (Slurm, KAI, Volcano, Kueue), GPU/InfiniBand/RDMA, storage architecture for AI/ML
Technologies
Go, Python, Kubernetes, Cilium, Multus, InfiniBand, RoCE, RDMA, Prometheus, Grafana, NVIDIA GPU stack
Responsibilities
Design and build scalable control plane services and custom Kubernetes controllers; develop automation for cluster lifecycle management; build GPU-aware orchestration systems; partner on networking solutions for AI workloads; write resilient systems for failure handling; develop platform services for inference and model serving; build internal tools and CLIs for ML/AI teams; support and debug production issues
Seniority
Senior, hands-on IC