Lead Site Reliability Engineer
Core
Lead a team of Site Reliability Engineers to build resilient, automated cloud infrastructure and ensure high availability for Glean's Work AI platform.
Role type
Senior IC Site Reliability Engineering Lead
Builds
Hybrid cloud infrastructure, automated production environments, and scalable AI agent systems for enterprise customers
Domain
Enterprise AI / Workforce Intelligence / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Cloud architecture (GCP/AWS/Azure), Kubernetes, Docker, Terraform, incident management, SLO definition, automation scripting, distributed systems design, security compliance
Preferred skills
Cross-team collaboration, architectural decision making, on-call process optimization
Technologies
Google Cloud Platform, AWS, Azure, Docker, Kubernetes, Terraform
Responsibilities
Lead technical strategy and mentorship for the SRE team; manage complex cloud scale challenges; participate in on-call rotation and drive blameless postmortems; develop automation tools for deployment and monitoring; optimize infrastructure for performance and cost; collaborate on security and compliance; provide SRE insights during system design reviews
Seniority
Senior, hands-on IC with team leadership