Senior Manager, Site Reliability Engineering
Core
Ensure availability, performance, and reliability of containerized business applications by bridging development and operations through troubleshooting, automation, and deployment management.
Role type
Senior Manager, Site Reliability Engineering
Builds
Containerized business applications for Lending, Payments, and Universal Banking
Domain
Financial Services / Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Container operations (Docker, Kubernetes), Bash scripting, Python scripting, Java application troubleshooting, Log analysis (Grafana, Loki), Linux system administration, Disaster recovery management
Preferred skills
CI/CD tools (Jenkins, GitHub CI, ArgoCD), SQL (Oracle), Cloud platforms (Azure, AWS, GCP), Container orchestration (AKS, Flux, Ansible), Messaging systems (Kafka, IBM MQ, Redhat AMQ), Observability tools (ElasticSearch, Grafana)
Responsibilities
Create and manage DevOps pipelines for application deployment; Diagnose complex runtime errors and performance bottlenecks in containerized/non-containerized apps; Respond to production alerts and lead root-cause analysis; Manage change records and audit processes; Validate application deployments across staging and production; Write scripts to automate health checks, log rotation, and recovery procedures; Configure monitoring dashboards and alerts
Seniority
Senior, hands-on IC with management responsibilities