Senior Technical Program Manager, Infrastructure
Core
Lead large-scale, cross-functional initiatives to define, scale, and optimize Glean's infrastructure platform, ensuring it remains performant, scalable, and resilient for enterprise AI workloads.
Role type
Senior Infrastructure Technical Program Manager (TPM)
Builds
Infrastructure platform powering Glean's intelligent Search, AI Assistant, and scalable AI agents for enterprise customers.
Domain
Enterprise AI / Cloud Infrastructure / Workforce Productivity
Deliverable
production ML models | infrastructure
Required skills
Technical program management, infrastructure engineering, reliability engineering (SRE), cloud infrastructure (AWS/GCP/Azure), distributed systems, data pipelines, ML training workflows, LLM runtime infrastructure, capacity planning, observability, cost optimization, deployment orchestration, stakeholder management, risk management, automation, cross-functional leadership
Preferred skills
Experience with ML/AI teams, understanding of agentic AI systems, experience in B2B/enterprise environments
Technologies
AWS, GCP, Azure, Kubernetes, Terraform, Prometheus, Grafana, MLflow, Ray, LangChain, Docker, Jenkins, GitLab CI, Jira, Confluence, Slack, Microsoft Teams, Zoom, ServiceNow, Zendesk, GitHub
Responsibilities
Drive the company's Infrastructure roadmap across Setup & Deployment, Runtime, Storage, and AI Infra; Lead cross-functional programs that improve scalability, reliability, cost efficiency, and developer velocity; Define and orchestrate how Glean instances are deployed, upgraded, and monitored at scale; Partner with AI and Data teams to evolve ML pipelines, model training infrastructure, and LLM serving stack; Lead initiatives to improve observability, configuration management, and resource utilization; Coordinate capacity planning, infrastructure migrations, and performance optimization programs; Build clear visibility into infra cost drivers and partner with finance and engineering leaders on optimization initiatives; Lead end-to-end infra programs spanning compute, networking, storage, orchestration, and AI workloads; Partner with Engineering to define standards for environment provisioning, deployment automation, and configuration governance; Develop and operationalize frameworks for runtime health, scaling, and disaster recovery; Drive consistency and automation across deployment orchestration systems; Establish clear metrics for reliability, performance, and cost efficiency; Coordinate cross-team delivery of high-impact programs such as data pipeline scalability, LLM infrastructure expansion, or infra observability improvements; Communicate program status and technical risks effectively to leadership and stakeholders; Continuously identify process or system bottlenecks, and drive automation to improve speed and reliability of infra operations
Seniority
Senior, hands-on IC with strategic oversight