Lead Site Reliability Engineer
Core
Lead SRE owning design, reliability, and evolution of a multi-cloud, multi-region content serving platform handling 25B+ daily requests.
Role type
Lead Site Reliability Engineer (IC + Strategy)
Builds
Multi-cloud content serving platform, logging platform, automation tooling, observability stack
Domain
Internet / Content Personalization / Multi-cloud Infrastructure
Deliverable
production ML models | product features | infrastructure
Required skills
Distributed systems architecture, Multi-cloud strategy (AWS/GCP), Infrastructure as Code, Kubernetes cluster design, Observability platform leadership, On-call excellence, Linux systems diagnosis, Automation strategy, Mentoring
Preferred skills
None explicitly stated as preferred
Technologies
Apache Pulsar, Apache Kafka, Grafana Loki, ScyllaDB/Cassandra, Prometheus, Thanos, Grafana Alloy, Loki, Tempo, Terraform, Chef, EKS, GKE, NodeJS, Golang, Ruby, Python, shell scripting
Responsibilities
Define and drive automation strategy for infrastructure tooling; Own design, reliability, and evolution of core platform applications; Architect and lead logging platform strategy; Establish capacity planning and performance management frameworks; Lead cross-functional reliability initiatives; Demonstrate high autonomy in addressing systemic weaknesses
Seniority
Senior, hands-on IC with strategic leadership