Site Reliability Engineering Lead
Core
Drive technical direction of production reliability, observability, and incident management for a global live entertainment platform serving 300M+ users.
Role type
Senior Site Reliability Engineering Lead (hands-on IC with people leadership)
Builds
High-availability distributed systems for live event discovery and management
Domain
Live entertainment technology / Data-driven event platforms
Deliverable
production ML models | infrastructure
Required skills
Site Reliability Engineering, incident management, observability (SLIs/SLOs/Error Budgets), distributed systems operations, automation, CI/CD, scripting (Python/Go/Bash), software engineering
Preferred skills
Kubernetes, cloud platforms (AWS/GCP/Azure), OpenTelemetry, event-driven architectures (Kafka), AI-powered engineering tools
Technologies
Datadog, Grafana, Prometheus, Terraform, Python, Go, Bash, Kubernetes, AWS, GCP, Azure, Kafka, OpenTelemetry
Responsibilities
Lead SRE culture to reduce MTTD/MTTR; define and improve SLIs/SLOs/Error Budgets; lead critical incident response and RCA; improve monitoring, logging, tracing, and automation; partner with cross-functional teams on production readiness; own operational processes and on-call practices; lead, coach, and develop the Tech Support Engineering team
Seniority
Senior, hands-on IC with team leadership