CareerPlanSign in

Senior / Lead Site Reliability Engineer

Watford, England, United Kingdom💼 Full-time🗓 2026-07-28 → 2026-09-26

Core

Technical leadership for reliability engineering across the digital estate, ensuring high availability and resilience for customer-facing lottery systems during normal operations and peak events.

Role type

Senior/Lead Site Reliability Engineer

Builds

Instant-Win and Draw-based platforms, Player Identity & Protection systems, CMS, and Geolocation services

Domain

Gaming/Lottery industry, Cloud Infrastructure (AWS)

Deliverable

production ML models | infrastructure

Required skills

SRE practices (SLOs, SLIs, error budgets), Incident management, Cloud architecture (AWS ECS/EKS), Infrastructure as Code (Terraform), Observability (Splunk, CloudWatch, Grafana), Distributed systems troubleshooting, Python/Go programming

Preferred skills

Container platform migration (ECS to EKS), High-scale consumer platform experience, Real-time analytics tooling

Technologies

AWS, Terraform, Splunk, CloudWatch, Grafana, Quantum Metric, Kubernetes, ECS, Python, Go

Responsibilities

Define and govern SLOs/SLIs/error budgets; Lead incident response and post-mortems; Drive automation and platform maturity; Own capacity planning for high-concurrency events; Mentor engineers on SRE practices

Seniority

Senior, hands-on IC with leadership responsibilities

Sourced via workable · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.