CareerPlanSign in

Hardware Health

United States, Multiple Locations, Multiple Locations💼 Full-time🗓 2025-11-25 → 2026-09-26

Core

Design and develop next-generation hardware health monitoring and diagnostic frameworks for large GPU clusters, building predictive analytics pipelines to anticipate hardware degradation and systemic issues.

Role type

Senior IC hardware reliability and observability engineer

Builds

Predictive analytics pipelines, real-time observability platforms, and automated health management systems for large-scale GPU clusters

Domain

AI hardware infrastructure, large-scale GPU clusters, datacenter operations

Deliverable

production ML models | infrastructure

Required skills

GPU architecture expertise, high-speed interconnects (NVLink, InfiniBand, RoCE), hardware telemetry and diagnostics, failure analysis, reliability modeling, predictive maintenance, automation development, C/C++/Python/Java programming

Preferred skills

Experience with exascale-class systems, cloud-scale AI clusters, machine learning-based anomaly detection, thermal efficiency optimization

Technologies

NVIDIA H100/GB200, NVLink, InfiniBand, RoCE, C, C++, C#, Java, JavaScript, Python

Responsibilities

Design and develop hardware health monitoring frameworks, build predictive analytics pipelines, lead incident triage for high-impact hardware issues, define system health KPIs, drive automation in health management, partner with cross-functional teams to influence hardware design

Seniority

Senior, hands-on IC

Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.