Principal Technical Program Management, AI Infrastructure
Core
Build and scale AI infrastructure stacks to host AI Foundry Services, enabling the transition of AI workloads from research to production.
Role type
Principal Technical Program Manager (AI Infrastructure)
Builds
Scalable distributed computing systems, multi-node GPU infrastructure, and AI inference/training fleets.
Domain
Cloud Infrastructure / AI Engineering
Deliverable
production ML models | infrastructure
Required skills
Technical program management, distributed platform ecosystem design, GPU/VM/OS/cloud fundamentals, cross-functional project leadership, product roadmap ownership, data-driven goal setting, Python for ML, cloud provider expertise (Azure/AWS/GCP), ML platform experience.
Preferred skills
Experience shipping complex products for developers/ML professionals, navigating ambiguity, influencing cross-functional teams.
Technologies
Azure, AWS, Google Cloud, Python, ML platforms, GPUs, VMs, OS.
Responsibilities
Define product requirements and manage end-to-end development for AI infrastructure; translate business goals into technical strategy; oversee experiments and measure success with data; collaborate with UX, Data Science, and Engineering teams to build consumer-grade applications.
Seniority
Principal, hands-on IC with strategic oversight