CareerPlanGet AI match score →

Site Reliability Engineer

Manchester, England, UK💼 Full-time🗓 2026-05-28 → 2026-05-31

Required skills

Experience of Linux administration, Experience or strong understanding of Kubernetes, Being comfortable in a scripting language suitable for automation tasks, Understanding of current recovery solutions and high availability architectures for cloud and on prem, Understanding of Capacity Management & Planning scenarios and tooling, Experience with Agile principles and practices, Expertise in problem diagnosis across complex, distributed systems

Preferred skills

Experience supporting SaaS products, Experience using AI for Automation, Experience with Incident Management, Post Mortems and related practices, Knowledge of observability and monitoring best practices, Experience operating within one or more public clouds (AWS, GCP, Azure), Experience with configuration management, and infrastructure as code

Technologies

Linux, Kubernetes, scripting language, cloud and on prem recovery solutions, high availability architectures, Capacity Management & Planning tooling, Agile, observability and monitoring, public clouds (AWS, GCP, Azure), configuration management, infrastructure as code

Responsibilities

Work with Engineering & Service Management to ensure that the disaster recovery and Capacity plans drive disaster recovery (DR) strategy and procedures both in Cloud and DC venues, Build out tooling that supports the DR plans and tracks progress and maturity against set KPI’s and Metrics, Work with Engineering & Service Management to ensure that disaster recovery solutions are adequate, in place, maintained, and tested as part of the regular operational life cycle, Provide ongoing feedback for risk management, mitigation, and prevention, Develop and implement capacity planning tooling, frameworks, policies, and strategies, Provide capacity requirements and impact assessments for new services or changes, Collaborate with other Platform managers to deliver objectives on our platform evolution roadmap

Seniority

Domain

Platform disaster response/crisis management, Capacity planning, Cloud and on prem infrastructure, SaaS products, Incident Management, observability and monitoring, public clouds

Sourced via linkedin · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on LinkedIn ↗