Site Reliability Engineer
Core
Advanced-level SRE responsible for establishing SRE practices, improving reliability/observability, and ensuring stability of critical banking technology services.
Role type
Senior Site Reliability Engineer
Builds
Resilient, highly available customer-facing banking services and automated incident response capabilities
Domain
Banking / Cloud Infrastructure (Azure)
Deliverable
production ML models | infrastructure | product features
Required skills
Advanced programming, SRE practices (SLOs/SLIs/SLAs), observability, incident response, release engineering, automation, disaster recovery, code review
Preferred skills
Azure cloud services, chaos engineering, self-healing systems, Blue-Green/Canary deployments, synthetic monitoring
Technologies
Azure Monitor, Dynatrace, source code management tools, synthetic monitoring platforms
Responsibilities
Define and mature SLOs/SLIs/SLAs, standardize observability dashboards, implement resiliency patterns (circuit breakers, autoscaling), improve incident response playbooks, support disaster recovery validation, conduct code reviews, partner with dev teams on release safety
Seniority
Senior, hands-on IC