Senior Manager, Site Reliability and Production Delivery Management (US)
Core
Lead reliability, availability, production deployment, and operational excellence practices across critical technology platforms to ensure highly available, secure, and resilient customer-facing applications.
Role type
Senior Manager, Site Reliability and Production Delivery Management
Builds
Enterprise standards for monitoring, incident management, release engineering, deployment automation, and service reliability.
Domain
Banking / Financial Services / Cloud Infrastructure / SRE
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
SRE practices, Observability (SLOs/SLIs/Error Budgets), Release Governance, Change Management, Incident Management, Team Leadership, Operational Metrics, Cloud Platforms, CI/CD, Automation
Preferred skills
Large enterprise environment leadership, SRE/Platform Engineering experience, ITIL processes, Mission-critical application management, Reliability transformation
Technologies
Dynatrace, Splunk, Datadog, Grafana, AppDynamics, Prometheus, ELK Stack, OpenTelemetry, Azure, AWS, Google Cloud Platform, Kubernetes, OpenShift, Docker, GitHub, Azure DevOps, Jenkins, Harness, ArgoCD, Python, PowerShell, Shell Scripting, Terraform, Ansible, ServiceNow
Responsibilities
Lead and mature platform Observability and SRE practices; Establish standards for dashboards, logging, tracing, incident response, and on-call support; Drive continuous improvement in reliability, resiliency, and automation; Lead release planning, deployment governance, and change management; Oversee production deployments, readiness reviews, and rollback strategies; Provide leadership during major incidents and post-incident reviews; Build and lead high-performing teams for reliability engineering and deployment governance.
Seniority
Senior Manager, strategic direction & team leadership