Lead Site Reliability Engineer
Core
Lead the reliability engineering practice for LSEG's internal observability platform, defining SLO/SLI frameworks, error budget policies, and incident response leadership to improve operational reliability.
Role type
Lead Site Reliability Engineer (IC/Lead)
Builds
Internal observability platform (telemetry, dashboards, alerts, service health)
Domain
Financial services infrastructure / Observability
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
SRE practice design, SLO/SLI framework design, distributed systems failure modes, incident response leadership, cross-team influence, mentoring SRE practitioners
Preferred skills
OpenTelemetry, Grafana, ClickHouse, Cribl, Datadog, BigPanda, Redis, PromQL, ClickHouse SQL, greenfield platform establishment, regulated financial services environment
Technologies
OpenTelemetry, Grafana, ClickHouse, Cribl, Datadog, BigPanda, Redis, PromQL
Responsibilities
Define and own the SLO and SLI framework, set error budget policy, build and improve the on-call programme, lead incident response and post-incident reviews, mentor SRE practitioners, provide reporting on reliability posture
Seniority
Lead, hands-on IC with leadership scope