CareerPlanSign in

Service Engineer

United States, Washington, Redmond💼 Full-time🗓 2026-08-26 → 2026-09-26

Core

Lead Azure's Incident Management practice, acting as the single point of command during high-severity incidents to restore services and protect customer trust.

Role type

Senior Incident Commander / Service Engineer (Cloud Operations)

Builds

Azure cloud infrastructure services and incident response frameworks

Domain

Cloud Operations / Site Reliability Engineering (SRE)

Deliverable

production ML models | product features | dashboards & analysis | client delivery | infrastructure

Required skills

Incident response and crisis management, cloud architecture patterns, microservices, containerization, monitoring and observability, automation scripting, ITIL frameworks, high availability and disaster recovery, root cause analysis, cross-functional leadership

Preferred skills

AI/ML integration in cloud infrastructure, chaos engineering, Windows/Linux debugging, cloud certifications (AWS/Azure/GCP), ITIL/SRE certifications

Technologies

Azure, AWS, GCP, Grafana, Prometheus, Datadog, Splunk, New Relic, PowerShell, Python

Responsibilities

Lead and manage high-severity incidents across Azure services as the central authority; drive incident reviews (RCAs/PIRs) and implement preventative improvements; collaborate with engineering and product teams to design resilient architecture; participate in on-call rotation; analyze telemetry to identify root causes; advocate for customer self-service capabilities and operational frameworks.

Seniority

Senior, hands-on IC with strategic oversight

Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.