Manager, Site Reliability Engineering
Core
Lead a high-performing SRE team to ensure the reliability, performance, security, and compliance of shared services platforms underpinning critical authentication, API, data intelligence, and data warehouse capabilities.
Role type
Manager, Site Reliability Engineering (SRE)
Builds
Shared Services platforms (Risk Intelligence Hub, Hyper Gateway, Lakehouse) supporting global Risk Intelligence offerings
Domain
Financial services / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Cloud-native operations on AWS, Kubernetes, Docker, CI/CD pipelines, Infrastructure as Code (Terraform), Observability platforms, Incident management, People leadership, Strategic planning
Preferred skills
Snowflake platform reliability, ITIL/ISO 27001 frameworks, Cloud financial management, Financial services domain experience
Technologies
AWS, Kubernetes, Docker, Terraform, GitHub Actions, Jenkins, Datadog, BigPanda, OpenTelemetry, Snowflake
Responsibilities
Own 24×7 reliability and operational excellence for Shared Services platforms; Define and improve SLIs, SLOs, and error budgets; Lead major incident response as Incident Commander; Champion automation-first approaches for incident response and deployment; Lead, mentor, and develop a team of SRE engineers; Partner with HR for workforce planning and strategic hiring in Bangalore; Represent the Bangalore site in global SRE forums; Govern vendor and partner engagements.
Seniority
Manager, hands-on technical and people leader