Technical Program Manager, Compute Infrastructure
Core
End-to-end delivery of large-scale GPU clusters and compute infrastructure supporting production AI models and training workloads.
Role type
Senior Technical Program Manager (Compute Infrastructure)
Builds
Production-ready GPU clusters, network Points-of-Presence (PoPs), and unified compute platforms for Applied AI and Research.
Domain
Hyperscale AI infrastructure, GPU fleet management, data center operations.
Deliverable
production ML models | infrastructure
Required skills
Program management for capital projects, cross-functional team leadership, vendor/partner ecosystem management, risk management, dependency management, executive reporting, technical partnership with engineering teams, supply chain logistics.
Preferred skills
Hard science degree, engineering expertise, experience with hyperscaler infrastructure deployment, experience interfacing with chip providers and construction firms.
Technologies
GPU clusters, networking hardware, power and cooling systems, rack cabling infrastructure.
Responsibilities
Lead end-to-end delivery of new compute SKUs and large-scale GPU clusters across external partners; drive multi-threaded bring-up programs spanning hardware, networking, power, and cooling; interface with chip providers to derisk new hardware onboarding; build and operationalize program mechanisms for predictable delivery; partner with engineering to improve cluster reliability and automation; support network operations and physical/logical bring-up of PoPs; coordinate cross-functional readiness for production shipping; manage integration and handoffs across teams and partners; identify bottlenecks and drive systemic fixes; provide executive visibility on portfolio progress and risks.
Seniority
Senior, hands-on IC