Senior Site Reliability Engineer
Core
Own the availability, performance, and scalability of critical shared services at LSEG by maintaining SLOs, writing automation, and ensuring rapid recovery.
Role type
Senior Site Reliability Engineer (IC)
Builds
Scalable, reliable, and performant cloud services and infrastructure
Domain
Financial markets infrastructure / Cloud Operations
Deliverable
production ML models | infrastructure
Required skills
Scripting (Shell, Python), Infrastructure as Code (Terraform), Cloud platforms (Azure/AWS), Container orchestration (Docker, Kubernetes), Observability (Datadog, Dynatrace), CI/CD pipelines, Incident response, Git workflows
Preferred skills
DevOps methodologies, Algorithms and data structures, Identity and access management, Relational databases (Azure)
Technologies
Terraform, Datadog, BigPanda, EntraID, Azure, AWS, Docker, Kubernetes, Shell, Python
Responsibilities
Maintain Service Level Objectives (SLOs) for owned systems; Write automation to scale systems and prevent/recover from service issues; Partner with development teams to improve reliability and release velocity; Participate in on-call rotations, incident response, and root cause analysis; Perform architectural reviews and operational testing for cloud migration projects; Configure observability dashboards and metrics.
Seniority
Senior, hands-on IC