Manager of Site Reliability Engineering (SRE)
Core
Lead and develop a team of SRE practitioners to deliver highly reliable, scalable, and performant cloud-based infrastructure and services.
Role type
Manager of Site Reliability Engineering
Builds
Cloud-based infrastructure and services on Google Cloud Platform (GCP)
Domain
Cloud Infrastructure / Site Reliability Engineering
Deliverable
infrastructure
Required skills
Site Reliability Engineering (SRE) principles, infrastructure-as-code (Terraform, ArgoCD), CI/CD pipelines, Kubernetes, container orchestration, cloud-native architecture, observability (Dynatrace, Datadog), incident management, capacity planning, system design
Preferred skills
Experience in large enterprise environments (Fortune 500), hybrid cloud integrations, API-driven/microservices architectures
Technologies
Google Cloud Platform (GCP), Kubernetes, Terraform, ArgoCD, Dynatrace, Datadog, Azure DevOps
Responsibilities
Lead and mentor SRE team, implement SRE principles and DevOps best practices, define and track SRE metrics, drive automation efforts (CI/CD, IaC), improve observability practices, participate in incident response and root cause analysis, partner with engineering/product/security teams, oversee cloud infrastructure management, develop long-term cloud platform roadmap, establish and monitor SLIs/SLOs
Seniority
Manager, team leadership & strategy