Site Reliability Engineer
Core
Ensure stability and reliability of the Reservations suite of products, focusing on automated delivery, incident response, and reducing toil.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Production services for the Reservations suite, supporting airlines, hoteliers, and travel agencies
Domain
Travel technology, Cloud-native infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Google Cloud Platform (GCP), Linux/UNIX, Terraform, shell scripting, networking, CI/CD, container orchestration (Docker, Kubernetes), monitoring and alerting (AppDynamics, Google Cloud ops, Data dog, Prometheus, Grafana, Elastic search), incident management, change management
Preferred skills
Java development, relational databases (Oracle, Couchbase, Datastore, Spanner), Google SRE practices, Service Now, runbooks
Technologies
GCP, Jenkins, Docker, Kubernetes, Terraform, AppDynamics, Data dog, Prometheus, Grafana, Elastic search, Oracle, Couchbase, Datastore, Spanner
Responsibilities
Provide application on-call support and troubleshoot system alerts; Lead investigation of severe incidents affecting customer experience; Build alerting and monitoring solutions; Engage in SRE practices like blameless postmortems and SLI building; Support development partners with infrastructure deployments and CI/CD pipelines; Manage infrastructure currency, PCI audits, capacity monitoring, and cost optimization; Provide technical mentorship to teams