Director Operational Resilience Engineering
Core
Lead root cause analyses and reliability engineering investigations into high-severity datacenter incidents to identify systemic failure mechanisms and drive permanent corrective actions.
Role type
Director, Operational Resilience Engineering
Builds
Fleet-wide reliability models, predictive metrics, and dashboards for global datacenter operations
Domain
Datacenter infrastructure, reliability engineering, forensic analysis
Deliverable
production ML models | dashboards & analysis | infrastructure
Required skills
Root cause analysis (RCA), Failure Modes and Effects Analysis (FMEA), Fault Tree Analysis (FTA), Weibull analysis, probabilistic risk assessment, statistical analysis, trend analysis, predictive analytics, telemetry integration, cross-functional team leadership, global program management, forensic engineering frameworks, risk assessment methodologies, corrective action planning, data interpretation
Preferred skills
AI-driven tools integration, multinational environment experience, industry standards knowledge (ISO, IEC)
Technologies
Telemetry systems, predictive analytics platforms, AI-driven forensic tools
Responsibilities
Lead RCA and reliability investigations for high-severity incidents, develop fleet-wide reliability models, embed findings into global operational standards, conduct systemic risk assessments, develop predictive metrics and dashboards, partner with Engineering and Operations teams, measure and validate corrective action effectiveness, establish best practices and continuous improvement methodologies
Seniority
Director, strategic leadership with hands-on technical expertise