CareerPlanGet AI match score →

Hardware Health

United States, Multiple Locations, Multiple Locations💼 Full-time🗓 2025-11-25 → 2026-07-31

Core

Design and develop next-generation hardware health monitoring and diagnostic frameworks for large GPU clusters, building predictive analytics pipelines to anticipate hardware degradation and systemic issues.

Role type

Senior IC hardware reliability and observability engineer

Builds

Predictive analytics pipelines, real-time observability platforms, and automated health management systems for large-scale GPU clusters

Domain

AI hardware infrastructure, large-scale GPU clusters, datacenter operations

Deliverable

production ML models | infrastructure

Required skills

GPU architecture expertise, high-speed interconnects (NVLink, InfiniBand, RoCE), hardware telemetry and diagnostics, failure analysis, reliability modeling, predictive maintenance, automation development, C/C++/Python/Java programming

Preferred skills

Experience with exascale-class systems, cloud-scale AI clusters, machine learning-based anomaly detection, thermal efficiency optimization

Technologies

NVIDIA H100/GB200, NVLink, InfiniBand, RoCE, C, C++, C#, Java, JavaScript, Python

Responsibilities

Design and develop hardware health monitoring frameworks, build predictive analytics pipelines, lead incident triage for high-impact hardware issues, define system health KPIs, drive automation in health management, partner with cross-functional teams to influence hardware design

Seniority

Senior, hands-on IC

Rewrite
## About the role Design and develop next-generation hardware health monitoring and diagnostic frameworks for large GPU clusters (NVL16/NVL72/GB200+ scale). Build predictive analytics pipelines leveraging telemetry, power, and thermal data to anticipate hardware degradation and systemic issues. Collaborate with silicon, firmware, and datacenter engineers to identify root causes and remediate large-scale hardware anomalies. Define system health KPIs (e.g., NIS/RIS, MTBF, failure domain analysis) and integrate them into real-time observability platforms. Lead incident triage for high-impact GPU, network, and cooling issues across distributed clusters. Drive automation in health management to reduce manual intervention to the top 5% of anomalies. Partner with cross-functional teams to influence hardware design for reliability, thermal efficiency, and serviceability. ## Requirements - Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python - Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python - OR Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python - OR equivalent experience. - Experience working with large-scale HPC or GPU systems (NVIDIA H100/GB200 or equivalent). - Deep understanding of GPU architecture, high-speed interconnects (NVLink, InfiniBand, RoCE), and large datacenter topologies. - Proficiency in hardware telemetry, diagnostics, or failure analysis tools. - Experience with exascale-class systems or cloud-scale AI clusters. - Familiarity with reliability modeling, machine learning-based anomaly detection, or predictive maintenance. - Contributions to large-scale infrastructure operations, supercomputing centers, or AI hardware design. ## Nice to have - Experience working with large-scale HPC or GPU systems (NVIDIA H100/GB200 or equivalent). - Deep understanding of GPU architecture, high-speed interconnects (NVLink, InfiniBand, RoCE), and large datacenter topologies. - Proficiency in hardware telemetry, diagnostics, or failure analysis tools. - Experience with exascale-class systems or cloud-scale AI clusters. - Familiarity with reliability modeling, machine learning-based anomaly detection, or predictive maintenance. - Contributions to large-scale infrastructure operations, supercomputing centers, or AI hardware design. ## What we offer - Competitive salary and equity package - Comprehensive health, dental, and vision benefits - 401(k) matching - Flexible work arrangements - Professional development opportunities - Collaborative and innovative work environment ## About us We are a leading technology company dedicated to advancing the future of computing. Our mission is to build the next generation of hardware and software solutions that empower organizations to solve the world's most complex challenges. Join us in shaping the future of AI, high-performance computing, and datacenter infrastructure.
Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply at Microsoft ↗