Principal Engineer, Cluster Orchestration
Core
Design and evolve cluster orchestration systems (Slurm, Kubernetes, SUNK) for large-scale AI training, inference, and model onboarding.
Role type
Principal Engineer, AI Infrastructure (Cluster Orchestration)
Builds
Kubernetes-native control planes, schedulers, admission logic, and internal tooling for GPU clusters.
Domain
AI Infrastructure / Cloud Computing / Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Kubernetes internals, Slurm internals, Go, distributed systems architecture, scheduling algorithms, multi-tenant GPU isolation, SLO definition, incident response
Preferred skills
Kueue, Kubeflow, Argo Workflows, Ray, Istio, Knative, ML platform engineering, open-source contributions
Technologies
Kubernetes, Slurm, SUNK, Kueue, Go
Responsibilities
Define long-term architecture for orchestration platforms; Lead evolution of control planes and custom operators; Set standards for reliability and observability; Write and review production code for controllers and schedulers; Mentor senior engineers and influence cross-functional teams
Seniority
Principal, strategy & mentorship