Manager Site Reliability Engineering
Core
Lead a team of SRE engineers to ensure the reliability, scalability, and operational excellence of Sabre Web Services, a critical API gateway and connectivity infrastructure for the travel industry.
Role type
Manager, Site Reliability Engineering
Builds
Highly available, fault-tolerant, cloud-native services for Sabre Web Services (API gateway and connectivity infrastructure)
Domain
Travel technology, Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
People leadership, Cloud engineering, Distributed systems, Observability, Incident management, Automation, AI-assisted operations
Preferred skills
Kubernetes, Terraform, CI/CD pipelines, SLOs, Error budgets, High-availability SaaS platforms
Technologies
Google Cloud Platform (GCP), Kubernetes, Infrastructure as Code, CI/CD automation, Observability platforms
Responsibilities
Lead and develop a team of Site Reliability Engineers, Drive service reliability and operational excellence across critical production systems, Partner with Engineering and Product teams to establish Service Level Objectives (SLOs), Lead major incident response and postmortem reviews, Promote automation and AI-enabled operations to reduce operational toil
Seniority
Manager, hands-on leadership