Tech Lead, Site Reliability Engineering
Core
Lead 24/7 reliability operations for critical Risk Screening applications (WC1 and World Check verify) supporting the EMEA region, ensuring uptime, performance, and compliance.
Role type
Senior IC/Lead Site Reliability Engineer (SRE)
Builds
Cloud-native Screening services for financial risk intelligence
Domain
Financial markets infrastructure / Cloud Operations
Deliverable
production ML models | product features | dashboards & analysis
Required skills
SRE principles (SLOs, SLIs, error budgets), cloud operations (Azure, AWS), container orchestration (Kubernetes, Docker), CI/CD, infrastructure-as-code (Terraform), observability, identity/fraud platform experience
Preferred skills
People management, strategic thinking, incident command
Technologies
Azure (SQL, Cosmos DB, Key Vault, Sentinel), AWS (Lambda, ECS, RDS, CloudWatch), Kubernetes, Docker, Terraform, GitHub Actions, Jenkins, Datadog, BigPanda, OpenTelemetry
Responsibilities
Define and maintain SLOs/SLIs; act as Major Incident Commander; drive automation for incident response and deployment; mentor SRE engineers; standardize observability and ITSM tooling; ensure DR readiness
Seniority
Senior, hands-on IC with people management