Sr. Technical Product Manager GPU Orchestration
Core
Define and deliver the roadmap for managed Kubernetes, managed Slurm services, and GPU orchestration platforms to support AI/ML training and inference workloads.
Role type
Senior Technical Product Manager (GPU Orchestration & Cloud Infrastructure)
Builds
Managed Kubernetes clusters, Slurm services, and GPU scheduling platforms for enterprise customers
Domain
Cloud Infrastructure, High-Performance Computing (HPC), AI/ML Infrastructure
Deliverable
production ML models | product features
Required skills
Product management in cloud infrastructure, Kubernetes, Slurm, GPU scheduling, resource allocation, multi-tenant isolation, API architecture, distributed systems, cluster lifecycle management, AI/ML infrastructure
Preferred skills
Experience with SUNK and Run:ai integration, designing self-service workflows, defining service-level objectives for control plane reliability
Technologies
Kubernetes, Slurm, SUNK, Run:ai, AI APIs
Responsibilities
Define and deliver the roadmap for managed Kubernetes and Slurm services; own the full cluster lifecycle including provisioning and scaling; build scheduling capabilities for GPU workloads; coordinate integration with networking and storage; design APIs and CLI tools for self-service administration
Seniority
Senior, hands-on IC