Lead Site Reliability Engineer
Core
Lead reliability engineering for critical identity and fraud screening applications supporting the EMEA region, ensuring availability, scalability, and security of business-critical platforms.
Role type
Senior IC/Lead Site Reliability Engineer (People Management)
Builds
High-availability Screening services (WC1 and World Check verify Applications) for financial markets infrastructure
Domain
Financial services / Identity & Fraud / Cloud Operations
Deliverable
production ML models | infrastructure
Required skills
SRE principles (SLOs, SLIs, error budgets), cloud operations (Azure, AWS), container orchestration (Kubernetes, Docker), CI/CD pipelines, infrastructure-as-code (Terraform), observability platforms, incident management, team leadership
Preferred skills
Experience with identity platforms and fraud detection systems, vendor management, offshore strategy planning
Technologies
Azure (SQL, Cosmos DB, Application Gateway, Key Vault, Storage, DNS, Load Balancer, VMs, Azure ML, Sentinel), AWS (Lambda, ECS, RDS, CloudWatch), Kubernetes, Docker, Terraform, GitHub Actions, Jenkins, Datadog, BigPanda, OpenTelemetry
Responsibilities
Lead 24×7 reliability operations and act as Major Incident Commander; define SLOs/SLIs and drive improvements via engineering backlogs; mentor SRE engineers and manage team performance; standardize tooling for observability and incident communication; ensure DR readiness and resilience patterns; partner with product teams to embed non-functional requirements
Seniority
Senior, hands-on IC with people management