Senior System Reliability Engineer
Core
Drive Design for Reliability (DfR) and Reliability, Availability & Maintainability (RAM) processes for Cloud & AI hardware and infrastructure solutions.
Role type
Senior System Reliability Engineer
Builds
Cloud & AI hardware and infrastructure solutions
Domain
Hardware reliability, Cloud infrastructure, AI systems
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Reliability Block Diagram (RBD) modeling, Markov modeling, Discrete Event Simulations (DES), Fault Tree Analysis (FTA), Prognostics & Health Management (PHM) modeling, Remaining Useful Life (RUL) prediction, DFMEA, Reliasoft, JMP, Python scripting, repairable/non-repairable system statistics
Preferred skills
Experience with Networking, Power, and/or Cooling infrastructure
Responsibilities
Set reliability and availability targets for sub-systems and components, develop tests to demonstrate reliability targets, carry out DFMEA to identify and mitigate critical risks, utilize modeling methods to quantify risks, develop PHM models for RUL prediction based on telemetry data
Seniority
Senior, hands-on IC