CareerPlanSign in

Director Operational Resilience Engineering

United States, Washington, Redmond💼 Full-time🗓 2026-07-08 → 2026-09-25

Core

Lead root cause analyses and reliability engineering investigations into high-severity datacenter incidents to identify systemic failure mechanisms and drive permanent corrective actions.

Role type

Director, Operational Resilience Engineering

Builds

Fleet-wide reliability models, predictive metrics, and dashboards for global datacenter operations

Domain

Datacenter infrastructure, reliability engineering, forensic analysis

Deliverable

production ML models | dashboards & analysis | infrastructure

Required skills

Root cause analysis (RCA), Failure Modes and Effects Analysis (FMEA), Fault Tree Analysis (FTA), Weibull analysis, probabilistic risk assessment, statistical analysis, trend analysis, predictive analytics, telemetry integration, cross-functional team leadership, global program management, forensic engineering frameworks, risk assessment methodologies, corrective action planning, data interpretation

Preferred skills

AI-driven tools integration, multinational environment experience, industry standards knowledge (ISO, IEC)

Technologies

Telemetry systems, predictive analytics platforms, AI-driven forensic tools

Responsibilities

Lead RCA and reliability investigations for high-severity incidents, develop fleet-wide reliability models, embed findings into global operational standards, conduct systemic risk assessments, develop predictive metrics and dashboards, partner with Engineering and Operations teams, measure and validate corrective action effectiveness, establish best practices and continuous improvement methodologies

Seniority

Director, strategic leadership with hands-on technical expertise

Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.