CareerPlanGet AI match score →

Senior System Reliability Engineer

United States, Washington, Redmond💼 Full-time🗓 2026-07-14 → 2026-07-20

Core

Drive Design for Reliability (DfR) and Reliability, Availability & Maintainability (RAM) processes for Cloud & AI hardware and infrastructure solutions.

Role type

Senior System Reliability Engineer

Builds

Cloud & AI hardware and infrastructure solutions

Domain

Hardware reliability, Cloud infrastructure, AI systems

Deliverable

production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work

Required skills

Reliability Block Diagram (RBD) modeling, Markov modeling, Discrete Event Simulations (DES), Fault Tree Analysis (FTA), Prognostics & Health Management (PHM) modeling, Remaining Useful Life (RUL) prediction, DFMEA, Reliasoft, JMP, Python scripting, repairable/non-repairable system statistics

Preferred skills

Experience with Networking, Power, and/or Cooling infrastructure

Responsibilities

Set reliability and availability targets for sub-systems and components, develop tests to demonstrate reliability targets, carry out DFMEA to identify and mitigate critical risks, utilize modeling methods to quantify risks, develop PHM models for RUL prediction based on telemetry data

Seniority

Senior, hands-on IC

Rewrite
## About the role Ability to drive both Design for Reliability (DfR) and Reliability, Availability & Maintainability (RAM) processes with consistency and rigor for both current and next generation Cloud & AI hardware and infrastructure solutions. Work across functionally across different disciplines (Hardware, Firmware, Architecture, Safety, Serviceability etc.) to influence business and operational decisions. Works with key stakeholders to set appropriate reliability and availability targets and allocates the targets to lower-level sub-systems and components. Develops appropriate tests at system, sub-system and/or/component as needed to demonstrate the budgeted reliability targets. Carries out effective Design Failure Modes & Effects Analysis (DFMEA) to identify and mitigate critical risks via design, operational and diagnostic improvements. Utilizes different modeling methods such as Reliability Block Diagram (RBD), Markov, Discreet Event Simulations (DES) and Fault Tree Analysis (FTA) to quantify risks and/or support business decisions. Develops Prognostics & Health Management (PHM) models for Remaining Useful Life (RUL) prediction based on telemetry data. ## Requirements - Doctorate Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 2+ years technical engineering experience - OR Master's Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 4+ years technical engineering experience - OR Bachelor's Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 5+ years technical engineering experience - OR 12+ years relevant technical engineering experience - M.S. in Electrical or Electronic Engineering, Reliability Engineering, or equivalent discipline 5+ years of experience performing Reliability Engineering on Complex Systems - Background in Applied Reliability Engineering statistics for repairable and non-repairable systems - Experienced in using software tools such as Reliasoft, JMP and python scripts for performing different reliability analyses - Passionate individual who is able to work collaboratively in a team environment and across internal divisions, industry (OEM, ODM), and with customers - Prior experience driving reliability efforts for Cloud & AI hardware including infrastructure (Networking, Power and/or Cooling) - Experienced in Prognostics & Health Management (PHM) techniques - Certified Reliability Engineer (CRE)
Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply at Microsoft ↗