Sr. Hardware Reliability Engineer, Infrastructure Reliability & Quality
Core
Proactively drive reliability risk identification, assessment, and mitigation for datacenter infrastructure equipment (e.g., Air Handling Units, Generators, Transformers, Chillers) and perform root cause analysis of critical failures to ensure datacenter availability.
Role type
Senior Hardware Reliability Engineer (Infrastructure)
Builds
AWS global datacenter infrastructure reliability and quality standards
Domain
Data Center Infrastructure / Hardware Reliability Engineering
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Physics-of-Failure approach, lifecycle environmental and operational stress analysis, statistical modeling, reliability block diagrams, vendor qualification, factory and site testing oversight, DFMEA support, end-of-life strategy development
Preferred skills
Proactive reliability approaches across design/manufacture/deployment, external supply chain partner management, cost-effective reliability optimization
Technologies
Finite Element Analysis (FEA), accelerated life testing tools, reliability quantification software
Responsibilities
Drive DFR methodology for new product designs; Qualify third-party critical infrastructure equipment; Oversee factory and site testing of equipment; Guide Root Cause Analysis of field failures; Analyze internal reliability data to create metrics; Provide feedback to procurement on vendor performance; Develop end-of-life strategies for critical equipment
Seniority
Senior, hands-on IC with program management