Senior Incident Manager
Core
Lead end-to-end management of high-visibility technical incidents and enterprise customer escalations, ensuring rapid restoration of services and effective communication throughout the lifecycle.
Role type
Senior Incident Manager (SRE/Reliability) (via careerplan.io/jobs/25ffe604-641f-49e7-a0f1-d2df3661a4fd-senior-incident-manager-at-crusoe)
Builds
AI infrastructure platform reliability and customer trust
Domain
AI Infrastructure / Cloud Services / Data Centers
Required skills
Incident response leadership, Cross-functional coordination, Customer communication, Root cause analysis (RCA), Incident metrics reporting, Process design, Linux familiarity, Kubernetes familiarity, Virtualization knowledge, Data analysis, Written communication
Preferred skills
Startup environment experience, Cloud infrastructure background, AI-focused environment experience, Slurm knowledge, TCP/IP stack understanding, Infrastructure-as-Code (IaC) practices
Technologies
Linux, Kubernetes, Slurm, TCP/IP, IaC
Responsibilities
Lead incident response and coordination across teams, Own status page communications and updates, Produce customer-facing root cause analyses, Track and report incident metrics, Design and implement incident response strategies, Participate in on-call rotation
Seniority
Senior, hands-on IC
