CareerPlanSign in

MTS - Site Reliability Engineer

United States, Washington, Redmond💼 Full-time🗓 2026-03-05 → 2026-09-26

Core

Ensure uptime, resiliency, and fault tolerance of AI model training and inference systems while building automation for hybrid cloud/on-prem environments.

Role type

Senior Site Reliability Engineer (AI Infrastructure)

Builds

Production-ready AI model training and inference pipelines on hybrid cloud/on-prem CPU+GPU infrastructure

Domain

Generative AI, Cloud Infrastructure, High-Performance Computing

Deliverable

production ML models

Required skills

Kubernetes, Docker, CI/CD pipelines, public cloud platforms (Azure/AWS/GCP), infrastructure-as-code, monitoring & observability tools (Grafana, Datadog, OpenTelemetry), Python/Go/Bash scripting, distributed systems, networking, storage

Preferred skills

Large-scale GPU cluster management, ML training/inference pipelines, HPC workload schedulers, capacity planning, cost optimization for GPU-heavy environments

Technologies

Kubernetes, Docker, Azure, AWS, GCP, Grafana, Datadog, OpenTelemetry

Responsibilities

Design and maintain monitoring, alerting, and logging systems for model serving pipelines; Analyze system performance and optimize resource utilization (compute, GPU clusters, storage, networking); Build automation for deployments, incident response, scaling, and failover; Lead on-call rotations and conduct blameless postmortems; Partner with ML engineers to accelerate research-to-production workflows

Seniority

Senior, hands-on IC

Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.