Senior / Lead Site Reliability Engineer
Core
Technical leadership for reliability engineering across the digital estate, ensuring high availability and resilience for customer-facing lottery systems during normal operations and peak events.
Role type
Senior/Lead Site Reliability Engineer
Builds
Instant-Win and Draw-based platforms, Player Identity & Protection systems, CMS, and Geolocation services
Domain
Gaming/Lottery industry, Cloud Infrastructure (AWS)
Deliverable
production ML models | infrastructure
Required skills
SRE practices (SLOs, SLIs, error budgets), Incident management, Cloud architecture (AWS ECS/EKS), Infrastructure as Code (Terraform), Observability (Splunk, CloudWatch, Grafana), Distributed systems troubleshooting, Python/Go programming
Preferred skills
Container platform migration (ECS to EKS), High-scale consumer platform experience, Real-time analytics tooling
Technologies
AWS, Terraform, Splunk, CloudWatch, Grafana, Quantum Metric, Kubernetes, ECS, Python, Go
Responsibilities
Define and govern SLOs/SLIs/error budgets; Lead incident response and post-mortems; Drive automation and platform maturity; Own capacity planning for high-concurrency events; Mentor engineers on SRE practices
Seniority
Senior, hands-on IC with leadership responsibilities