Director of Incident Response
Core
Building and leading the full incident response function for high-performance computing (HPC) and multi-tenant cloud environments, managing major incidents, forensic readiness, and engineering remediation loops.
Role type
Director of Incident Response (Builder)
Builds
End-to-end IR function including staffing, playbooks, tooling, forensic capabilities, and executive reporting for HPC and cloud infrastructure.
Domain
High-Performance Computing (HPC) and Multi-tenant Cloud Security
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Incident command for Sev-0/1 events, forensic readiness (bare-metal, Kubernetes, hypervisor), detection-to-containment runbook development, root cause analysis with engineering remediation, team leadership and hiring, regulatory notification management, metric-driven performance tracking (MTTD/MTTR), Jira Service Management, PagerDuty, SIEM/XDR platforms, NIST SP 800-61r2, MITRE ATT&CK, CIS Controls.
Preferred skills
GCIH/GCFA/GCFR certifications, ITIL 4, experience in regulated environments (financial services, export-controlled), CSP independence or infrastructure repatriation experience.
Technologies
Slurm, InfiniBand, GPU clusters, Kubernetes, AWS, Azure, GCP, Jira Service Management, PagerDuty, SIEM, XDR.
Responsibilities
Stand up IR function from ground up (staffing, rotations, tooling), command major incidents, develop detection-to-containment runbooks, establish forensic readiness, drive root cause analysis to engineering remediation, maintain Known Error Database and tabletop exercises, instrument IR with hard metrics, partner on detection engineering feedback loops, own executive/board reporting, co-own business continuity testing.
Seniority
Director, hands-on builder & leader