Service Engineer II
Core
Lead incident response and drive platform reliability improvements for Azure's global cloud infrastructure.
Role type
Senior IC Service Engineer (Incident Command & SRE)
Builds
High-availability cloud services and automated incident response systems
Domain
Cloud Infrastructure / Site Reliability Engineering
Deliverable
production ML models | infrastructure
Required skills
Incident command and crisis management, SRE principles, cloud architecture patterns, microservices, containerization, automation scripting, observability tooling, root cause analysis, ITIL frameworks, high availability design, disaster recovery, business continuity, performance tuning, strategic thinking, quantitative analysis, team leadership, cross-functional collaboration, problem resolution, judgment, decision-making, customer communication, task prioritization, debugging, AI/ML integration understanding, chaos engineering, fault injection
Preferred skills
Windows/Linux platform expertise, developer tools, AI-powered incident analysis, predictive alerting, cloud certifications (AWS/Azure/GCP), ITIL/SRE certifications, customer compassion, agility, technical excellence
Technologies
Azure, AWS, GCP, Grafana, Prometheus, Datadog, Splunk, New Relic, PowerShell, Python, CLI
Responsibilities
Lead and manage high-severity incidents as the single point of accountability; drive post-incident reviews and process development; identify and implement platform-wide improvements using telemetry; design next-generation architecture for cloud infrastructure services; collaborate with engineering and product teams to implement service resiliency enhancements; participate in on-call rotation; analyze customer-impacting signals to identify root causes and drive preventative improvements; advocate for customer self-service capabilities and scalable solutions; design and drive adoption of incident response playbooks and operational frameworks.
Seniority
Senior, hands-on IC with strategic leadership