Team Lead Software Systems Engineering
Core
Ensuring stability and reliability of the Reservations suite of products through SRE practices, automated delivery, and incident management.
Role type
Lead Site Reliability Engineer (SRE)
Builds
Production-ready infrastructure, CI/CD pipelines, and observability solutions for the Reservations suite
Domain
Travel technology, Cloud-native systems
Deliverable
infrastructure
Required skills
Google Cloud Platform, Linux/UNIX, Terraform, shell scripting, networking, CI/CD concepts, Jenkins, Docker, Kubernetes, monitoring and alerting tools (AppDynamics, Google Cloud ops, Data dog, Prometheus, Grafana, Elastic search), analysis and debugging, change management
Preferred skills
Java development, relational databases (Oracle, Couchbase, Datastore, Spanner), Google SRE knowledge, Service Now, runbooks
Technologies
GCP, Jenkins, Docker, Kubernetes, Terraform, AppDynamics, Data dog, Prometheus, Grafana, Elastic search, Oracle, Couchbase, Datastore, Spanner
Responsibilities
Provide application on-call support and troubleshoot system alerts, lead investigation of severe incidents, build alerting and monitoring solutions, continuously improve reliability via SRE practices, support development partners with infrastructure deployments, take ownership of services including capacity monitoring and cost optimization, provide technical mentorship
Seniority
Lead, hands-on IC with mentorship