Sr. Incident Management Engineer
Core
Lead mitigation and resolution of high-severity service and infrastructure incidents while coordinating cross-functional engineering teams during live-site events.
Role type
Senior Incident Management Engineer
Builds
Reliable infrastructure and services for customers
Domain
Cloud operations and Site Reliability Engineering
Deliverable
production ML models | infrastructure
Required skills
Incident management, cloud operations, SRE, infrastructure engineering, complex technical incident leadership, troubleshooting, analytical skills, automation scripting, operational process improvement, playbook development, escalation path design, outage preparedness, HW/EE/thermal/mechanical issue identification
Preferred skills
Data-driven decision-making, stakeholder influence, cross-functional collaboration
Technologies
KQL, Azure Monitor, Power BI
Responsibilities
Lead mitigation and resolution of high-severity service and infrastructure incidents; Coordinate cross-functional engineering teams during live-site events; Drive clear communication and decision-making during incident response; Ensure timely escalation, mitigation, and customer impact reduction; Improve incident management processes, response playbooks, and escalation paths; Support operational readiness reviews and outage preparedness activities
Seniority
Senior, hands-on IC