CareerPlanGet AI match score →

Senior Site Reliability Engineer

United States, Washington, Redmond💼 Full-time🗓 2026-07-21 → 2026-07-31

Core

Design and implement automation solutions to improve service reliability, availability, and performance for enterprise customers while providing operational insights to product teams.

Role type

Senior Site Reliability Engineer (Customer Support & Reliability)

Builds

Automation tooling, proactive alerting systems, and service telemetry enhancements for large-scale distributed systems.

Domain

Cloud Infrastructure & Enterprise Support

Deliverable

production ML models | infrastructure

Required skills

Large-scale cloud services operations, Service Reliability/Availability/Performance improvement, Observability and MELT implementation, Logic Apps, Jupyter Notebooks, incident root cause analysis and automation, distributed systems troubleshooting, service telemetry design.

Preferred skills

None stated.

Technologies

Logic Apps, Jupyter Notebooks

Responsibilities

Collaborate with engineering teams to build automation for faster issue resolution; interface with enterprise customers for service escalations; design telemetry changes for automation consumption; analyze data to provide operational insights to product teams; influence product architecture for supportability.

Seniority

Senior, hands-on IC

Rewrite
## About the role - Collaborating closely with engineering teams on building and enhancing tooling and automation solutions for faster resolution of issues impacting SLO's and averting incidents altogether when possible. - Collaborating with the customers to understand their pain points around supportability and SLO attainment and formulate strategies for addressing recurring issues in a sustainable way. - Communicate on a deeply technical level and be the single point of contact for interfacing with enterprise customers for handling service escalations and driving the issues to resolution. - Ability to design and implement any changes to service telemetry for the automation to consume if it is not already available. - Enhancing customer facing experience by proactive alerting based on utilization, trends, resource health, etc. - Analyze data and provide operational insights into customer experience to design and product teams, so that we can design features with supportability in mind. - Embody our culture and values. ## Requirements - 6+ years technical experience in software engineering, network engineering, or systems administration OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 3+ years technical experience in software engineering, network engineering, or systems administration OR Master's Degree in Computer Science, Information Technology, or related field AND 2+ years technical experience in software engineering, network engineering, or systems administration. - These requirements include, but are not limited to the following specialized security screenings: - 4+ years of experience running large scale cloud services. - 2+ years of operational experience in improving Service Reliability, Availability and Performance. - Understanding of Observability and MELT implementation patterns for large-scale services. - Experience in Logic Apps and authoring Jupyter Notebooks. - Experience in analyzing, troubleshooting, and automating root cause analysis and mitigation of incidents impacting large-scale distributed systems. - Systematic problem-solving approach, coupled with effective communication skills and a sense of curiosity. - Ability to deal with the ambiguity associated with working in a fast-paced environment. - Influencing the product architecture and roadmap to make sure the customer-experienced supportability is always a key consideration when evolving the product.
Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply at Microsoft ↗