CareerPlanSign in

Product Manager, Inference Platform

San Francisco💼 Full-time🗓 2026-04-02 → 2026-09-26

Core

Define the product strategy and roadmap for a mission-critical inference platform that enables AI companies to deploy, scale, and manage large language models reliably across multi-cloud environments.

Role type

Senior Product Manager, Inference Platform

Builds

A unified platform for production AI inference, including autoscaling, traffic routing, failover, and release management.

Domain

Cloud Infrastructure / AI Model Serving

Deliverable

production ML models

Required skills

Product management for infrastructure/distributed systems, end-to-end ownership (backend to UX), cross-team roadmap driving, defining new categories, scaling/routing/failover reasoning

Preferred skills

GPU infrastructure, Kubernetes, serving frameworks (vLLM, TensorRT-LLM, SGLang)

Technologies

Kubernetes, vLLM, TensorRT-LLM, SGLang, Multi-cloud

Responsibilities

Own workload scaling policies (autoscaling, placement, compliance), ensure production reliability (traffic routing, failover, health recovery), build release engines (canary, shadow, A/B), drive cost/performance optimization, reduce MTTR via self-serve incident management, set the roadmap for infrastructure teams.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.