CareerPlanGet AI match score →

Director Operational Resilience Engineering

United States, Washington, Redmond💼 Full-time🗓 2026-07-08 → 2026-07-16

Core

Lead root cause analyses and reliability engineering investigations into high-severity datacenter incidents to identify systemic failure mechanisms and drive permanent corrective actions.

Role type

Director, Operational Resilience Engineering

Builds

Fleet-wide reliability models, predictive metrics, and dashboards for global datacenter operations

Domain

Datacenter infrastructure, reliability engineering, forensic analysis

Deliverable

production ML models | dashboards & analysis | infrastructure

Required skills

Root cause analysis (RCA), Failure Modes and Effects Analysis (FMEA), Fault Tree Analysis (FTA), Weibull analysis, probabilistic risk assessment, statistical analysis, trend analysis, predictive analytics, telemetry integration, cross-functional team leadership, global program management, forensic engineering frameworks, risk assessment methodologies, corrective action planning, data interpretation

Preferred skills

AI-driven tools integration, multinational environment experience, industry standards knowledge (ISO, IEC)

Technologies

Telemetry systems, predictive analytics platforms, AI-driven forensic tools

Responsibilities

Lead RCA and reliability investigations for high-severity incidents, develop fleet-wide reliability models, embed findings into global operational standards, conduct systemic risk assessments, develop predictive metrics and dashboards, partner with Engineering and Operations teams, measure and validate corrective action effectiveness, establish best practices and continuous improvement methodologies

Seniority

Director, strategic leadership with hands-on technical expertise

Rewrite
## About the role Lead root cause analyses (RCA) and reliability engineering investigations into high-severity datacenter incidents to identify systemic failure mechanisms and drive permanent corrective actions. Apply advanced reliability engineering methodologies, including Failure Modes and Effects Analysis (FMEA), Fault Tree Analysis (FTA), Weibull analysis, reliability modeling, and probabilistic risk assessment. Perform incident forensics, statistical analysis, and trend analysis to identify recurring failure patterns, emerging risks, and opportunities for systemic improvement. Develop and maintain fleet-wide reliability models that quantify operational risk, predict failure behavior, and prioritize mitigation efforts. Embed RCA findings and reliability engineering insights into global operational standards, engineering requirements, maintenance strategies, and operating procedures. Conduct systemic risk assessments across critical power, cooling, controls, and operational processes to improve resilience and availability. Develop predictive metrics, leading indicators, and reliability dashboards that enable proactive identification and mitigation of operational risks. Partner with Engineering, Regional Operations, Design, and Product teams to drive fleet-wide reliability improvements throughout the datacenter lifecycle. Measure and validate the effectiveness of corrective actions using data-driven analysis to ensure sustainable elimination of recurring failure mechanisms. Establish best practices, analytical frameworks, and continuous improvement methodologies that improve the reliability, availability, resilience, and scalability of the global datacenter fleet. ## Requirements - Doctorate Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 5+ years technical engineering experience - Master's Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 7+ years technical engineering experience - Bachelor's Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 8+ years technical engineering experience - These requirements include, but are not limited to the following specialized security screenings: - Doctorate Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 7+ years technical engineering experience - Master's Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 10+ years technical engineering experience - Bachelor's Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 12+ years technical engineering experience - 4+ years people management experience - Experience integrating telemetry, predictive analytics, and AI-driven tools into forensic processes - Prior experience in multinational environments - Professional certifications such as Certified Reliability Engineer (CRE), or equivalent - Proven track record in Root Cause Analysis (RCA) and post-incident governance - Knowledge of risk assessment methodologies and containment strategies - Familiarity with forensic frameworks, evidence handling, and reliability metrics - Experience leading cross-functional teams and managing global programs - Advanced capability in failure analysis, data interpretation, and corrective action planning - Deep understanding of industry standards for forensic engineering (e.g., ISO, IEC) ## Nice to have - Professional certifications such as Certified Reliability Engineer (CRE), or equivalent ## What we offer - Partner with Engineering, Regional Operations, Design, and Product teams to drive fleet-wide reliability improvements throughout the datacenter lifecycle ## About us - Embody our culture and values
Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply at Microsoft ↗