Senior Site Reliability Engineer
Core
Design, build, and maintain a multi-cloud, multi-region, active-active content serving platform handling 25+ billion daily requests.
Role type
Senior Site Reliability Engineer (hands-on IC)
Builds
Scalable content serving infrastructure and core applications
Domain
Cloud infrastructure, content delivery, multi-region systems
Deliverable
production ML models | infrastructure
Required skills
Cloud platform expertise (AWS/GCP), Infrastructure as Code (Terraform), Kubernetes (EKS/GKE), Linux, High-level programming (NodeJS, Go, Ruby, Python), Shell Scripting, Observability (Prometheus, Thanos, Grafana, Loki, Tempo), On-call incident management
Preferred skills
Experience architecting large-scale observability platforms, Defining SLO frameworks, Multi-tenancy strategies
Technologies
AWS, GCP, Terraform, Kubernetes, EKS, GKE, Prometheus, Thanos, Grafana Alloy, Loki, Tempo, NodeJS, Go, Ruby, Python, Shell Scripting
Responsibilities
Improve infrastructure tooling and automation to minimize manual work and reduce incidents; Build, maintain, and support core applications; Monitor systems for capacity and performance; Partner with the SRE team for service delivery; Anticipate and address systemic weaknesses in the platform
Seniority
Senior, hands-on IC
