Staff Software Engineer, Cluster Orchestration
Core
Design and operate the orchestration platform (SUNK/Slurm on Kubernetes) to enable AI training and inference at scale across massive GPU clusters.
Role type
Staff Software Engineer, Cluster Orchestration
Builds
Kubernetes-native orchestration platform (SUNK) and managed services for AI workloads
Domain
AI Infrastructure / Cloud Computing / Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Go, distributed systems design, Kubernetes internals, Slurm internals, cloud-native development, technical strategy, cross-team architecture
Preferred skills
Ray, Kubeflow, Kueue, Istio, Knative, Argo Workflows, GPU-based applications, ML pipelines, scheduling strategies (quota enforcement, pre-emption), reliability practices (SLOs, post-incident reviews), AI infrastructure (ML training, inference, HPC), mentorship
Technologies
Kubernetes, Slurm, Go, Ray, Kubeflow, Kueue, Istio, Knative, Argo Workflows
Responsibilities
Define architectural direction for the orchestration platform, own critical parts of the platform and managed services, drive cross-org initiatives in scheduling and scaling, mentor senior engineers, establish org-wide best practices in reliability and observability
Seniority
Staff, technical leadership & strategy