CareerPlanSign in

Site Reliability Engineer (HPC)

United States, Multiple Locations, Multiple Locations💼 Full-time🗓 2026-02-17 → 2026-09-26

Core

Ensure uptime, resiliency, and fault tolerance of HPC clusters powering MAI model training and inference.

Role type

Site Reliability Engineer (HPC)

Builds

HPC clusters for ML model training and inference

Domain

High-Performance Computing / Machine Learning Infrastructure

Deliverable

infrastructure

Required skills

Kubernetes, Docker, container orchestration, CI/CD pipelines, public cloud platforms (Azure/AWS/GCP), infrastructure-as-code, monitoring & observability tools (Grafana, Datadog, OpenTelemetry), Python, Go, Bash, distributed systems, networking, storage, GPU cluster management, workload schedulers, capacity planning

Preferred skills

ML training/inference pipelines, high-performance computing (HPC) and workload schedulers, background in capacity planning & cost optimization for GPU-heavy environments

Technologies

Kubernetes, Docker, Grafana, Datadog, OpenTelemetry, Azure, AWS, GCP

Responsibilities

Ensure uptime, resiliency, and fault tolerance of HPC clusters; Design and maintain monitoring, alerting, and logging systems; Build automation for deployments, incident response, scaling, and failover; Lead on-call rotations, troubleshoot production issues, and conduct blameless postmortems; Ensure data privacy, compliance, and secure operations; Partner with ML engineers and platform teams to improve developer experience

Seniority

Mid-level to Senior, hands-on IC

Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.