CareerPlanSign in

Member of Technical Staff (TPM, Inference)

San Francisco💼 Full-time🗓 2026-09-02 → 2026-09-26

Core

Orchestrates the inference platform roadmap, coordinating between model providers, engineering, and product teams to ensure smooth model onboarding, capacity scaling, and production reliability.

Role type

Senior technical program manager (inference platform)

Builds

High-throughput inference stack serving Ask, Computer, and API traffic for first-party and third-party models

Domain

AI/ML infrastructure, distributed systems, model serving

Deliverable

production ML models

Required skills

Technical program management, production LLM inference, cross-functional orchestration, data and metrics analysis, infrastructure systems, cost-efficiency optimization, release management, vendor coordination

Preferred skills

Experience with GPU capacity planning, operating model design for new functions, agile team leadership

Technologies

LLM inference stacks, distributed systems, GPU compute infrastructure

Responsibilities

Execute the inference platform roadmap including request handling, rate limits, and usage controls; Coordinate model provider onboarding and launch readiness; Drive latency, throughput, uptime, and cost-efficiency metrics; Run operating models for model-release and optimization programs; Lead cross-functional delivery for inference-stack changes; Build release mechanisms like rituals and dashboards to ensure low-risk deployments.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.