Head of Platform Product Reliability
Core
Lead system-level product reliability for AI servers, accelerator platforms, rack systems, and datacenter infrastructure from design through field deployment.
Role type
Head of Platform Product Reliability (Hardware Systems)
Builds
AI inference hardware platforms (servers, racks, datacenter infrastructure)
Domain
AI Infrastructure / Hardware Systems / Datacenter
Deliverable
production ML models | product features | physical/clinical work
Required skills
System-level reliability engineering, failure analysis (root cause), reliability modeling (MTBF, Weibull, FIT), qualification planning (HALT/HASS/ALT/AST), cross-functional leadership, technical trade-off judgment
Preferred skills
Liquid-cooled systems, high-density power delivery, hyperscale datacenter deployment, fleet telemetry analytics, ODM/JDM partner management
Technologies
HALT, HASS, ALT, AST, Weibull analysis, FMEA, MTBF, FIT rate, thermal cycling, power cycling
Responsibilities
Define end-to-end reliability strategy and standards; Establish qualification methodologies and validation gates; Lead root-cause investigations for hardware failures; Develop system reliability models and projections; Build fleet reliability infrastructure (telemetry, monitoring); Drive product release readiness reviews; Build and lead the reliability engineering organization
Seniority
Senior, hands-on IC with leadership responsibilities