Principal Product Manager, AI Infrastructure and Orchestration
Core
Own the control plane layer for deploying, scaling, and governing AI agents and models on Kubernetes and heterogeneous accelerators for regulated, air-gapped, and sovereign enterprise deployments.
Role type
Principal Product Manager, AI Infrastructure and Orchestration
Builds
The deployment and workload API, placement/capacity logic, autoscaling policies, agent runtime isolation, traffic routing, governance/audit, and metering for inference workloads.
Domain
AI Infrastructure, Cloud Platforms, Kubernetes, GPU Accelerators, Sovereign/Regulated Deployments
Deliverable
production ML models | product features | infrastructure
Required skills
Product management for infrastructure/cloud services, Kubernetes architecture (scheduler, CRDs, operators, RBAC), GPU/accelerator behavior and topology-aware placement, multi-tenancy and isolation models, API product design and contract management, technical writing and prototyping
Preferred skills
Service networking (ingress, load balancing, TLS, private links), modern serving stacks (vLLM, KV cache, quantization), agentic workload patterns (session affinity, tool-call fan-out), air-gapped/sovereign deployment experience, CNCF contributions
Technologies
Kubernetes, Kubernetes API server, Device plugins, vLLM, Kubernetes schedulers, CRDs, Operators, RBAC, GPU accelerators
Responsibilities
Define API contracts and lifecycle semantics for workload deployment; Design placement and capacity allocation strategies across heterogeneous nodes; Establish autoscaling signals and cost-versus-latency trade-off policies; Define agent runtime isolation and state management; Design traffic routing and tenancy boundaries; Establish governance, audit trails, and access control policies; Define metering, quota, and pricing models for inference workloads.
Seniority
Principal, hands-on IC with strategic influence