Lead Site Reliability Engineer
Core
Building a Site Reliability Engineering center to own reliability and operational maturity of batch-critical settlement platforms processing credit and debit transactions.
Role type
Manager-level Lead Site Reliability Engineer
Builds
Settlement cycles, observability dashboards, and automation for batch processing systems
Domain
Financial services / Payments / High-availability batch processing
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
SRE, production operations, reliability engineering, Java, Python, Go, Cloud Native technologies, container orchestration, Shell/Bash scripting, Unix/Linux system administration
Preferred skills
Agentic AI automation tools, troubleshooting distributed systems, payments/financial services domain knowledge, Networking concepts
Technologies
Java, Python, Shell, SQL, AWS, Kubernetes, OpenShift, Datadog, Observe, HashiCorp Vault, agentic AI tools
Responsibilities
Own reliability for batch settlement systems, build and improve observability for settlement pipelines, drive automation of operational toil, partner with UK-based settlement engineers, participate in incident management, contribute to regulatory readiness
Seniority
Manager-level, hands-on IC with team ownership