Staff, Reliability Engineer
Core
Define reliability strategy for next-generation AI computing systems across data center and workstation products, ensuring high uptime and durability.
Role type
Staff Reliability Engineer (AI Hardware)
Builds
High-performance AI platforms (silicon, thermal, mechanical, electrical, software)
Domain
AI Hardware / High-Performance Computing
Deliverable
production ML models | product features | infrastructure
Required skills
Reliability engineering strategy, statistical analysis (HALT, HASS, ALT, MTBF, Weibull, FMEA), root-cause analysis, advanced cooling validation (vapor chambers, heat pipes, liquid cooling), cross-functional leadership, NPI process management, mentorship
Preferred skills
Experience in high-performance computing, AI hardware, or data center systems
Technologies
RISC-V, thermal lab equipment, liquid cooling systems
Responsibilities
Set reliability strategy for AI computing systems, lead root-cause investigations on failures, validate advanced cooling systems, act as point of contact for manufacturing partners, mentor engineers and lead design reviews
Seniority
Staff, strategic leadership with hands-on technical execution