Vice President, Enterprise Technology Command Center, Major Incident Manager, DTI Site Reliability Engineering
Core
End-to-end management of critical and major incidents impacting DBS's technology services, focusing on minimizing service disruption and driving continuous improvement in incident response processes.
Role type
VP, Major Incident Manager for Enterprise Technology Command Center (SRE Operations)
Builds
Operational resilience and stability across critical applications and infrastructure for a major bank
Domain
Financial Services / Banking / Site Reliability Engineering
Deliverable
client delivery
Required skills
Major incident management, cross-functional team leadership, SRE principles, regulatory compliance (MAS), blameless incident reviews, trend analysis, SLO/SLI definition, automation strategy, stakeholder communication, team mentorship
Preferred skills
Cloud platforms (AWS, Azure, GCP), incident management tools (ServiceNow, PagerDuty), monitoring tools (Grafana, Splunk), scripting (Python, Shell), GenAI proficiency, ITIL/SRE certifications
Technologies
AWS, Azure, GCP, ServiceNow, PagerDuty, Grafana, Splunk, Dynatrace, Elk, Python, Shell
Responsibilities
Lead and manage major incidents from detection through resolution; Act as primary communication point during incidents; Coordinate technical teams to diagnose and resolve complex production issues; Facilitate blameless incident reviews (BIRs); Analyze incident trends to identify systemic issues; Develop strategies to reduce MTTD, MTTA, and MTTR; Mentor junior incident managers and SRE operations staff
Seniority
VP, strategic leadership & hands-on management