Senior AI Scheduling & Orchestration Engineer
Core
Design and implement advanced batch scheduling architectures and cluster-wide admission control for high-volume AI workload traffic.
Role type
Senior IC distributed systems engineer (AI scheduling & orchestration)
Builds
Kubernetes-based scheduling stacks for multi-node gang scheduling and accelerator management
Domain
Cloud infrastructure + AI/ML distributed training
Deliverable
production ML models
Required skills
Kubernetes scheduling frameworks, distributed systems engineering, PyTorch Distributed, Ray, MPI, GPU hardware architectures, infrastructure automation, Terraform, Go-based Operators
Preferred skills
HPC environment operations, multi-tenancy isolation policies, NVLink/InfiniBand optimization
Technologies
Kubernetes, PyTorch Distributed, Ray, MPI, Terraform, Go, MIG, InfiniBand, NVLink
Responsibilities
Design multi-node gang scheduling architectures; Develop cluster-wide admission control and job queueing; Architect topology-aware pod placement for NVLink/InfiniBand; Implement GPU sharing (MIG, time-slicing) and multi-tenancy policies; Collaborate on bare-metal hardware and storage I/O integration; Drive scheduling-stack reliability and scalability; Mentor junior engineers and conduct design reviews
Seniority
Senior, hands-on IC