CareerPlanGet AI match score →

Site Reliability Engineer

United States, Multiple Locations, Multiple Locations💼 Full-time🗓 2026-07-15 → 2026-07-23

Core

Designated Responsible Individual (DRI) monitoring service health, managing on-call rotations, and implementing solutions for performance and functionality issues in large-scale distributed systems.

Role type

Senior Site Reliability Engineer (Infrastructure & Distributed Systems)

Builds

Production and deployment automation for complex product features; mitigations for Live Site service issues.

Domain

Cloud infrastructure, distributed systems, GPU/InfiniBand hardware support

Deliverable

production ML models | infrastructure

Required skills

On-call incident management, data collection and classification for system health, automation development, physical infrastructure management, large-scale cloud/distributed systems experience, project lifecycle ownership

Preferred skills

People management, broad influence and communication, end-to-end project management

Technologies

Cloud platforms, distributed systems, GPUs, InfiniBand

Responsibilities

Monitor service for degradation and downtime; collect and analyze system metrics; develop automation for production and deployment; implement solutions for complex performance issues; manage physical infrastructure supporting GPUs and InfiniBand.

Seniority

Senior, hands-on IC with people management experience

Rewrite
## About the role - Acts as a Designated Responsible Individual (DRI) working on call to monitor service for degradation, downtime, or interruptions. - Alerts stakeholders as to the status and gains approval to restore system/product/service for simple problems. - Contributes to efforts to collect, classify, and analyze data with little oversight on a range of metrics (e.g., health of the system, where bugs might be occurring). - Contributes to the refinement of product features by escalating findings from analyses to inform decisions regarding the engineering of products. - Contributes to the development of automation within production and deployment of a complex product feature. - Contributes to efforts to ensure the correct processes are followed to achieve a high degree of security, privacy, safety, and accessibility. - Implements solutions and mitigations to more complex issues impacting performance or functionality of Live Site service and escalates as necessary. ## Requirements - Master's Degree in Computer Science, Information Technology, or related field AND 1+ year(s) technical experience in software engineering, network engineering, or systems administration OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 2+ years technical experience in software engineering, network engineering, or systems administration OR equivalent experience. - 1+ years experience managing physical infrastructure. - Master's Degree in Computer Science, Information Technology, or related field AND 3+ years technical experience in software engineering, network engineering, or systems administration OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 5+ years technical experience in software engineering, network engineering, or systems administration OR equivalent experience. - 2+ years technical experience working with large-scale cloud or distributed systems. - 1+ year(s) people management experience. - Experience working on large-scale distributed services with on-call responsibilities. - Ability to build and influence broadly towards common goals and priorities. - Ownership of end-to-end project lifecycle with solid project management and communication skills. - Experience with managing physical infrastructure, supporting GPUs and InfiniBand. ## Nice to have - (None specified in the original text) ## What we offer - (None specified in the original text) ## About us - (None specified in the original text)
Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply at Microsoft ↗