Senior Software Engineer, Cluster Orchestration
Core
Building and scaling the orchestration platform (SUNK/Slurm on Kubernetes) to manage AI training and inference workloads across massive GPU clusters.
Role type
Senior IC Software Engineer (Cluster Orchestration)
Builds
SUNK (Slurm on Kubernetes) and Kubernetes-native foundation for AI compute
Domain
AI Cloud Infrastructure / Distributed Systems
Deliverable
production ML models
Required skills
Go, Kubernetes (production scale), distributed systems, observability (Prometheus/Grafana/OpenTelemetry), SLO/SLI definition, performance optimization
Preferred skills
Ray/Kubeflow/Kueue/Istio/Argo Workflows, GPU-based applications, scheduling concepts (quota/pre-emption/scaling), reliability practices (post-incident reviews)
Technologies
Kubernetes, Slurm, Prometheus, Grafana, OpenTelemetry, Go, Python, C++
Responsibilities
Own multiple services within the orchestration platform, lead design/code reviews, decompose projects into milestones, define SLIs/SLOs, mentor IC1/IC2 engineers
Seniority
Senior, hands-on IC