Site Reliability Engineering (SRE) Manager
Core
Leads SRE, Production Support, and Observability teams to ensure reliability, availability, and operational excellence of critical business applications and platforms.
Role type
Senior IC manager (SRE)
Builds
Resilient, scalable, and secure cloud-native and hybrid application environments
Domain
Financial services / Cloud infrastructure / Site Reliability Engineering
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
SRE principles, incident management, root cause analysis, automation, cloud platforms, distributed systems, stakeholder management, team leadership
Preferred skills
Azure cloud technologies, observability platforms (Dynatrace, Splunk, Datadog, Grafana), scripting (PowerShell, Python, Bash), CI/CD, Infrastructure as Code, operational AI use cases
Technologies
Azure, containers, APIs, microservices, Dynatrace, Splunk, Datadog, Grafana, Azure Monitor, OpenTelemetry, PowerShell, Python, Bash
Responsibilities
Define and execute SRE strategies, lead major incident response, establish observability standards, drive automation initiatives, recruit and develop high-performing teams, ensure adherence to enterprise risk and cybersecurity standards
Seniority
Senior, hands-on IC with management responsibilities