Staff Software Engineer, Cluster Orchestration
Core
Design and operate the core orchestration platform (SUNK/Slurm on Kubernetes) to manage AI training and inference workloads at scale across massive GPU clusters.
Role type
Staff Software Engineer, Cluster Orchestration
Builds
Kubernetes-native orchestration platform (SUNK) and managed services for AI workloads
Domain
Cloud infrastructure, AI/ML, Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Go, Kubernetes internals, Slurm internals, distributed systems design, cloud-native development, technical strategy, cross-team architecture
Preferred skills
Ray, Kubeflow, Kueue, Istio, Knative, Argo Workflows, GPU-based applications, ML pipelines, scheduling concepts (quota enforcement, pre-emption), reliability practices (SLOs, post-incident reviews), AI infrastructure
Responsibilities
Define architectural direction for the orchestration platform, own critical parts of the platform, drive cross-org initiatives in scheduling and scaling, mentor senior engineers, establish org-wide best practices in reliability and observability
Seniority
Staff, technical leader & strategy