CareerPlanGet AI match score →

Technical Program Manager, Inference Performance

San Francisco, CA💼 Full-time💰 $290,000–$290,000🗓 2026-07-14 → 2026-07-31

Core

Bridge between inference systems and the broader organization, driving strategic initiatives for inference runtime and accelerator performance.

Role type

Technical Program Manager, Inference Performance

Builds

Inference runtime and accelerator performance infrastructure for large language models

Domain

AI Infrastructure / Large Language Model Deployment

Deliverable

infrastructure

Required skills

Technical program management, Inference systems knowledge, Compiler or hardware accelerator understanding, Cross-functional coordination, Strategic planning, Stakeholder management, Process improvement

Preferred skills

Experience with infrastructure scaling, Hardware integrations, Deployment governance, Ambiguous environment navigation

Technologies

Inference runtime, Hardware accelerators, Large language models

Responsibilities

Lead cross-functional initiatives for new infrastructure integration, Partner with engineering teams to identify optimization opportunities, Drive end-to-end readiness for model and feature launches, Own and prioritize the inference deployment roadmap, Build strong relationships across research, engineering, and product teams, Identify inefficiencies in current workflows and drive systematic improvements

Seniority

Mid-to-Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗