Rewrite
## About the role
Lead root cause analyses (RCA) and reliability engineering investigations into high-severity datacenter incidents to identify systemic failure mechanisms and drive permanent corrective actions. Apply advanced reliability engineering methodologies, including Failure Modes and Effects Analysis (FMEA), Fault Tree Analysis (FTA), Weibull analysis, reliability modeling, and probabilistic risk assessment. Perform incident forensics, statistical analysis, and trend analysis to identify recurring failure patterns, emerging risks, and opportunities for systemic improvement. Develop and maintain fleet-wide reliability models that quantify operational risk, predict failure behavior, and prioritize mitigation efforts. Embed RCA findings and reliability engineering insights into global operational standards, engineering requirements, maintenance strategies, and operating procedures. Conduct systemic risk assessments across critical power, cooling, controls, and operational processes to improve resilience and availability. Develop predictive metrics, leading indicators, and reliability dashboards that enable proactive identification and mitigation of operational risks. Partner with Engineering, Regional Operations, Design, and Product teams to drive fleet-wide reliability improvements throughout the datacenter lifecycle. Measure and validate the effectiveness of corrective actions using data-driven analysis to ensure sustainable elimination of recurring failure mechanisms. Establish best practices, analytical frameworks, and continuous improvement methodologies that improve the reliability, availability, resilience, and scalability of the global datacenter fleet.
## Requirements
- Doctorate Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 5+ years technical engineering experience
- Master's Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 7+ years technical engineering experience
- Bachelor's Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 8+ years technical engineering experience
- These requirements include, but are not limited to the following specialized security screenings:
- Doctorate Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 7+ years technical engineering experience
- Master's Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 10+ years technical engineering experience
- Bachelor's Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 12+ years technical engineering experience
- 4+ years people management experience
- Experience integrating telemetry, predictive analytics, and AI-driven tools into forensic processes
- Prior experience in multinational environments
- Professional certifications such as Certified Reliability Engineer (CRE), or equivalent
- Proven track record in Root Cause Analysis (RCA) and post-incident governance
- Knowledge of risk assessment methodologies and containment strategies
- Familiarity with forensic frameworks, evidence handling, and reliability metrics
- Experience leading cross-functional teams and managing global programs
- Advanced capability in failure analysis, data interpretation, and corrective action planning
- Deep understanding of industry standards for forensic engineering (e.g., ISO, IEC)
## Nice to have
- Professional certifications such as Certified Reliability Engineer (CRE), or equivalent
## What we offer
- Partner with Engineering, Regional Operations, Design, and Product teams to drive fleet-wide reliability improvements throughout the datacenter lifecycle
## About us
- Embody our culture and values
Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.