Site Reliability Engineer
Core
On-call first responder for platform incidents, driving resolution end-to-end and reducing future incident frequency.
Role type
Senior Site Reliability Engineer
Builds
Resilient platform infrastructure, automated remediation tooling, and observability stacks for online gaming services.
Domain
iGaming / Online Gaming
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Incident triage and resolution, post-incident review leadership, observability stack ownership (alerting, dashboards, tracing), SLO/SLI definition and tracking, capacity planning, load testing, release reliability gating, infrastructure hardening, runbook creation, self-healing script development, automated remediation tooling, root cause analysis.
Preferred skills
Experience in iGaming, fintech, or other regulated 24/7 industries.
Technologies
Observability tools, distributed tracing systems, alerting platforms, log aggregation systems, load testing frameworks.
Responsibilities
Serve as on-call first responder for platform incidents; Lead triage and coordinate with DevOps and Engineering to drive resolution; Own post-incident reviews to close root causes; Own the full observability stack including alerting, dashboards, and distributed tracing; Define and enforce health standards for critical services; Continuously refine alert thresholds to improve signal quality; Systematically reduce incident frequency and blast radius; Analyse trends and track SLI/SLO performance; Drive platform improvements in configuration, architecture, and operational behaviour; Identify and eliminate manual, repetitive operational work; Build runbooks, self-healing scripts, and automated remediation tooling; Define and track Service Level Indicators and Objectives; Work upstream with Engineers to bake reliability into releases and capacity planning; Collaborate with DevOps on infrastructure hardening; Participate in release planning as a reliability gate.
Seniority
Senior, hands-on IC