Service Engineer
Core
Lead Azure's Incident Management practice, acting as the single point of command during high-severity incidents to restore services and protect customer trust.
Role type
Senior Incident Commander / Service Engineer (Cloud Operations)
Builds
Azure cloud infrastructure services and incident response frameworks
Domain
Cloud Operations / Site Reliability Engineering (SRE)
Deliverable
production ML models | product features | dashboards & analysis | client delivery | infrastructure
Required skills
Incident response and crisis management, cloud architecture patterns, microservices, containerization, monitoring and observability, automation scripting, ITIL frameworks, high availability and disaster recovery, root cause analysis, cross-functional leadership
Preferred skills
AI/ML integration in cloud infrastructure, chaos engineering, Windows/Linux debugging, cloud certifications (AWS/Azure/GCP), ITIL/SRE certifications
Technologies
Azure, AWS, GCP, Grafana, Prometheus, Datadog, Splunk, New Relic, PowerShell, Python
Responsibilities
Lead and manage high-severity incidents across Azure services as the central authority; drive incident reviews (RCAs/PIRs) and implement preventative improvements; collaborate with engineering and product teams to design resilient architecture; participate in on-call rotation; analyze telemetry to identify root causes; advocate for customer self-service capabilities and operational frameworks.
Seniority
Senior, hands-on IC with strategic oversight