Senior Incident Manager
Core
Lead end-to-end management of high-visibility technical incidents and enterprise customer escalations, ensuring rapid restoration of services and effective communication throughout the lifecycle.
Role type
Senior Incident Manager (SRE/Reliability)
Builds
AI infrastructure platform (compute, networking, storage, Kubernetes, Slurm)
Domain
AI Infrastructure / Cloud Services
Deliverable
production ML models | infrastructure (via careerplan.io/jobs/78c8be37-93e2-40dd-9be6-93ba06ec2613-senior-incident-manager-at-crusoe)
Required skills
Incident response leadership, cross-team coordination, customer-facing root cause analysis (RCA), incident metrics reporting, process design, on-call management, Linux familiarity, Kubernetes familiarity, TCP/IP stack understanding
Preferred skills
Startup environment experience, cloud infrastructure background, virtualization/orchestration knowledge, Infrastructure-as-Code (IaC) practices
Responsibilities
Lead incident response and coordination across teams, own status page communications and customer updates, produce customer-facing root cause analyses, track and report incident metrics, document corrective actions, develop training materials and knowledge base articles, design incident response strategies, participate in on-call rotation
Seniority
Senior, hands-on IC
