CareerPlanGet AI match score →

SRE Lead

Dublin, Ireland💼 Full-time🗓 2026-07-06 → 2026-07-31

Core

Leading the operational health and production readiness of a cloud-native, event-driven microservices payment platform migrating from a legacy monolith.

Role type

Senior SRE Lead (Site Reliability Engineering)

Builds

Cloud-native payment processing platform on AWS (Kubernetes, Confluent Cloud, Aurora PostgreSQL, Java/Quarkus)

Domain

Fintech / High-volume distributed systems

Deliverable

production ML models | product features | dashboards & analysis | infrastructure

Required skills

SRE practices, observability (OpenTelemetry, Prometheus, Dynatrace), SLI/SLO definition, resilience engineering (chaos testing, circuit breakers), AWS & Kubernetes, event-driven systems (Kafka), capacity planning, automation

Preferred skills

Payments domain experience, incident management tooling, chaos engineering frameworks

Technologies

AWS, Kubernetes, Confluent Cloud, Aurora PostgreSQL, Java, Quarkus, Dynatrace, Splunk, Fluent Bit, Prometheus, Argo Rollouts, OpenTelemetry

Responsibilities

Define production readiness standards (observability, SLOs, runbooks), build and maintain observability stack, define and implement SLIs/SLOs, own incident response practices, drive resilience engineering (chaos testing, load testing), support engineering teams on operability design, manage legacy-to-new system transition, automate operational toil, perform capacity planning

Seniority

Senior, hands-on IC with leadership responsibilities

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗