CareerPlanSign in

Site Reliability Engineer

London, UK💼 Full-time🗓 2026-03-18 → 2026-08-07

Core

Architecting and deploying observability platforms to monitor system health, performance, and reliability while driving AI-driven alerting and proactive anomaly detection.

Role type

Senior Site Reliability Engineer (SRE)

Builds

Autonomous operations, self-healing automation, and observability platforms

Domain

Cloud-native distributed systems & microservices

Deliverable

production ML models | infrastructure

Required skills

SRE principles, observability (Dynatrace, Datadog), Python scripting, Ansible, AWS, Azure, Docker, Kubernetes, CI/CD pipelines, chaos engineering, AI/ML for predictive analytics

Preferred skills

Large-scale SRE implementation, strategic mindset balancing engineering excellence with business priorities

Technologies

Datadog, Dynatrace, AWS, Azure, Python, Ansible, Docker, Kubernetes, Gremlin, Chaos Monkey

Responsibilities

Implement strategies for modernizing IT operations enhancing observability and toil reduction; Architect and deploy observability platforms to monitor system health, performance, and reliability effectively; Propose & drive strategies for AI-driven alerting and proactive anomaly detection to reduce MTTD & MTTR; Develop and enforce SRE best practices, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets; Establish & create AIOPS roadmap for improving operational efficiency; Lead efforts to automate repetitive tasks (toil) using scripting, orchestration tools, and AI/ML-based solutions; Drive incident management and root cause analysis processes through automation, ensuring continuous improvement to enable autonomous operations; Mentor and guide teams on adopting SRE principles and tools.

Seniority

Senior, hands-on IC with strategic mentorship

Sourced via adzuna · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.