Hardware Health
Core
Design and develop next-generation hardware health monitoring and diagnostic frameworks for large GPU clusters, building predictive analytics pipelines to anticipate hardware degradation and systemic issues.
Role type
Senior IC hardware reliability and observability engineer
Builds
Predictive analytics pipelines, real-time observability platforms, and automated health management systems for large-scale GPU clusters
Domain
AI hardware infrastructure, large-scale GPU clusters, datacenter operations
Deliverable
production ML models | infrastructure
Required skills
GPU architecture expertise, high-speed interconnects (NVLink, InfiniBand, RoCE), hardware telemetry and diagnostics, failure analysis, reliability modeling, predictive maintenance, automation development, C/C++/Python/Java programming
Preferred skills
Experience with exascale-class systems, cloud-scale AI clusters, machine learning-based anomaly detection, thermal efficiency optimization
Technologies
NVIDIA H100/GB200, NVLink, InfiniBand, RoCE, C, C++, C#, Java, JavaScript, Python
Responsibilities
Design and develop hardware health monitoring frameworks, build predictive analytics pipelines, lead incident triage for high-impact hardware issues, define system health KPIs, drive automation in health management, partner with cross-functional teams to influence hardware design
Seniority
Senior, hands-on IC