CareerPlanGet AI match score →

Principal Engineer, Cluster Orchestration

Bellevue, WA💼 Full-time💰 $206,000–$206,000🗓 2026-07-01 → 2026-07-31

Core

Design and evolve cluster orchestration systems (Slurm, Kubernetes, SUNK) for large-scale AI training, inference, and model onboarding.

Role type

Principal Engineer, AI Infrastructure (Cluster Orchestration)

Builds

Kubernetes-native control planes, schedulers, admission logic, and internal tooling for GPU clusters.

Domain

AI Infrastructure / Cloud Computing / Distributed Systems

Deliverable

production ML models | infrastructure

Required skills

Kubernetes internals, Slurm internals, Go, distributed systems architecture, scheduling algorithms, multi-tenant GPU isolation, SLO definition, incident response

Preferred skills

Kueue, Kubeflow, Argo Workflows, Ray, Istio, Knative, ML platform engineering, open-source contributions

Technologies

Kubernetes, Slurm, SUNK, Kueue, Go

Responsibilities

Define long-term architecture for orchestration platforms; Lead evolution of control planes and custom operators; Set standards for reliability and observability; Write and review production code for controllers and schedulers; Mentor senior engineers and influence cross-functional teams

Seniority

Principal, strategy & mentorship

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗