Senior Backend Engineer, Managed Kubernetes & Slurm
Core
Build control planes, backend services, and automation to provision, operate, and scale Kubernetes and Slurm clusters across a large-scale GPU fleet for AI training and inference.
Role type
Senior Backend Engineer (Managed Kubernetes & Slurm)
Builds
Production-ready GPU environments for model training, inference, and deployment
Domain
Cloud Infrastructure / AI Compute / Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Go or Python, Kubernetes, Slurm, distributed systems, cloud-native architectures, networking, storage, observability, CI/CD, incident response
Preferred skills
Kubernetes operators/controllers/CRDs, Slurm administration, Infrastructure-as-Code (Terraform, Crossplane, Helm, Argo CD, Flux, Kustomize), GPU infrastructure, event-driven systems (gRPC, messaging), open source contributions
Technologies
Go, Python, Kubernetes, Slurm, Terraform, Crossplane, Helm, Argo CD, Flux, Kustomize, gRPC
Responsibilities
Design and build backend services in Go or Python; Develop control plane services for cluster orchestration; Build distributed systems for automation and lifecycle management; Improve platform reliability, scalability, and security; Diagnose and resolve complex production issues; Collaborate with infrastructure and AI teams; Contribute to technical design, architecture, and on-call operations
Seniority
Senior, hands-on IC