CareerPlanSign in

Senior Incident Manager

💼 Full-time🗓 2026-06-03 → 2026-09-25

Core

Lead critical incident response for AI data center infrastructure, coordinating rapid resolution of service-impacting events across GPU clusters, networking, and operations.

Role type

Senior Incident Manager (Infrastructure Reliability)

Builds

High-availability AI cloud infrastructure serving researchers, enterprises, and hyperscalers

Domain

AI Cloud Infrastructure / Data Center Operations

Deliverable

production ML models | infrastructure

Required skills

Incident command & leadership, root cause analysis, cross-team coordination, crisis communication, operational decision making, infrastructure reliability, ITIL/SRE frameworks, incident tracking tools (PagerDuty, ServiceNow, Jira, Datadog, Prometheus/Grafana)

Preferred skills

AI/HPC infrastructure operations, SRE background, high-density GPU environment experience, hyperscale/colocation data center experience, Incident Command System (ICS) knowledge

Technologies

GPU clusters, InfiniBand networks, NVIDIA clusters, networking and storage infrastructure, cloud/hybrid platforms

Responsibilities

Lead SEV-1/SEV-2 incident response and serve as Incident Commander, coordinate engineering/networking/facilities/vendor teams during outages, conduct post-incident reviews and root cause analysis, track incident metrics (MTTR, MTTD), improve response processes and tooling, maintain incident documentation and playbooks, provide executive-level incident summaries

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.