Senior AI Engineer (Kubernetes & Customised Scheduler)
Core
Design, develop, and operate a proprietary Kubernetes-native job scheduler for massive-scale AI factories, enabling efficient execution of training, inference, and agentic workloads across GPU systems.
Role type
Senior IC platform engineer (Kubernetes scheduler)
Builds
Proprietary job scheduler and workload orchestration layer for AI factories
Domain
AI infrastructure / High-performance computing / Cloud platforms
Deliverable
production ML models
Required skills
Kubernetes, custom controllers, admission webhooks, scheduling plugins, topology-aware placement, AI-factory resource management, workload control policies, observability, API/CLI development
Preferred skills
Experience with Slurm, Slinky, Kueue, Volcano, YuniKorn, NVIDIA DSX OS concepts
Technologies
Kubernetes, GPU device plugins, RDMA, NVLink, NVSwitch, Slurm, Slinky, Kueue, Volcano, YuniKorn
Responsibilities
Design and build topology-aware placement mechanisms for GPU clusters; Implement workload-control capabilities including preemption and gang scheduling; Integrate scheduler with Kubernetes control-plane and observability systems; Develop APIs, SDKs, and CLI tools for workload management; Enable Model-to-Grid benchmarking by exposing scheduling data.
Seniority
Senior, hands-on IC (via careerplan.io/jobs/5417753008-senior-ai-engineer-kubernetes-customised-scheduler-at-firmus-technologies)