CareerPlanGet AI match score →

HPC Operations Engineering Manager

United States, Multiple Locations, Multiple Locations💼 Full-time🗓 2026-02-11 → 2026-07-31

Core

Lead a team of SREs to ensure uptime, resiliency, and fault tolerance of AI model training and inference systems in hybrid cloud/on-prem environments.

Role type

Senior HPC Operations Engineering Manager

Builds

Automated deployment, incident response, scaling, and failover systems for CPU+GPU environments

Domain

High-performance computing, AI/ML infrastructure, hybrid cloud

Deliverable

infrastructure

Required skills

Kubernetes, Docker, container orchestration, Python, Go, Bash, monitoring & observability tools (Grafana, Datadog, OpenTelemetry), CI/CD pipelines, distributed systems, networking, storage, capacity planning, cost optimization

Preferred skills

Large-scale GPU cluster management, ML training/inference pipelines, HPC workload schedulers, Kubernetes operators

Technologies

Azure, AWS, GCP, Kubernetes, Docker, Grafana, Datadog, OpenTelemetry

Responsibilities

Lead team of SREs, design monitoring/alerting/logging systems, build automation for deployments and incident response, manage on-call rotations and postmortems, ensure security and compliance, partner with ML engineers to improve developer experience

Seniority

Senior, hands-on IC with people management

Rewrite
## About the role - Team leadership: Lead a team of experienced SREs to ensure uptime, resiliency and fault tolerance of AI model training and inference systems. - Observability: Design and help maintain monitoring, alerting, and logging systems to provide real-time visibility into model serving pipelines and infra. - Automation & Tooling: Lead building of automation for deployments, incident response, scaling, and failover in hybrid cloud/on-prem CPU+GPU environments. - Incident Management: Lead on-call rotations, troubleshoot production issues, conduct blameless postmortems, and drive continuous improvements. - Security & Compliance: Ensure data privacy, compliance, and secure operations across model training and serving environments. - Collaboration: Partner with ML engineers and platform teams to improve developer experience and accelerate research-to-production workflows. ## Requirements - Bachelor's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with Site Reliability Engineering, DevOps, or Infrastructure Engineering Leadership roles AND 8+ years experience with Kubernetes, Docker, and container orchestration, AND 6+ years experience with programming/scripting skills not limited to Python, Go, or Bash - Master's Degree in Computer Science or related technical field AND 12+ years technical engineering experience AND 10+ years experience with Kubernetes, Docker, and container orchestration, AND 10+ years' experience with public cloud platforms like Azure/AWS/GCP and infrastructure-as-code OR equivalent experience - 6+ years people management experience. - 8+ years experience in monitoring & observability tools (Grafana, Datadog, OpenTelemetry, etc.). - Knowledge of CI/CD pipelines for Inference and ML model deployment. - Solid knowledge of distributed systems, networking, and storage. - Experience running large-scale GPU clusters for ML/AI workloads (preferred). - Familiarity with ML training/inference pipelines. - Experience with high-performance computing (HPC) and workload schedulers ( Kubernetes operators). - Background in capacity planning & cost optimization for GPU-heavy environments
Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply at Microsoft ↗