CareerPlanSign in

Site Reliability Engineer

United States, Multiple Locations, Multiple Locations💼 Full-time🗓 2026-07-15 → 2026-09-26

Core

Designated Responsible Individual (DRI) monitoring service health, managing on-call rotations, and implementing solutions for performance and functionality issues in large-scale distributed systems.

Role type

Senior Site Reliability Engineer (Infrastructure & Distributed Systems)

Builds

Production and deployment automation for complex product features; mitigations for Live Site service issues.

Domain

Cloud infrastructure, distributed systems, GPU/InfiniBand hardware support

Deliverable

production ML models | infrastructure

Required skills

On-call incident management, data collection and classification for system health, automation development, physical infrastructure management, large-scale cloud/distributed systems experience, project lifecycle ownership

Preferred skills

People management, broad influence and communication, end-to-end project management

Technologies

Cloud platforms, distributed systems, GPUs, InfiniBand

Responsibilities

Monitor service for degradation and downtime; collect and analyze system metrics; develop automation for production and deployment; implement solutions for complex performance issues; manage physical infrastructure supporting GPUs and InfiniBand.

Seniority

Senior, hands-on IC with people management experience

Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.