CareerPlanSign in

HPC Operations Engineering Manager

United States, Multiple Locations, Multiple Locations💼 Full-time🗓 2026-02-11 → 2026-09-26

Core

Lead a team of SREs to ensure uptime, resiliency, and fault tolerance of AI model training and inference systems in hybrid cloud/on-prem environments.

Role type

Senior HPC Operations Engineering Manager

Builds

Automated deployment, incident response, scaling, and failover systems for CPU+GPU environments

Domain

High-performance computing, AI/ML infrastructure, hybrid cloud

Deliverable

infrastructure

Required skills

Kubernetes, Docker, container orchestration, Python, Go, Bash, monitoring & observability tools (Grafana, Datadog, OpenTelemetry), CI/CD pipelines, distributed systems, networking, storage, capacity planning, cost optimization

Preferred skills

Large-scale GPU cluster management, ML training/inference pipelines, HPC workload schedulers, Kubernetes operators

Technologies

Azure, AWS, GCP, Kubernetes, Docker, Grafana, Datadog, OpenTelemetry

Responsibilities

Lead team of SREs, design monitoring/alerting/logging systems, build automation for deployments and incident response, manage on-call rotations and postmortems, ensure security and compliance, partner with ML engineers to improve developer experience

Seniority

Senior, hands-on IC with people management

Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.