CareerPlanGet AI match score →

Hardware Operations Engineer

💼 Full-time🗓 2026-06-20 → 2026-07-31

Core

Senior technical authority for hardware reliability and fleet health at large-scale AI datacenters, ensuring uptime for thousands of GPUs and servers.

Role type

Senior IC datacenter hardware operations lead

Builds

Reliable, high-performance AI compute infrastructure for frontier model training and inference

Domain

AI/ML infrastructure, hyperscale datacenter operations, GPU clusters

Deliverable

production ML models

Required skills

hardware failure diagnosis, root cause analysis, fleet health management, server platform expertise, rack integration, operational procedure development, spare parts strategy, cross-functional leadership, technical documentation

Preferred skills

GPU cluster support, fleet health telemetry, failure analysis methodologies (FRACAS, RCCA, FMEA), Linux system administration, NPI-to-sustaining transitions, EHS practices

Technologies

GPU systems, server platforms, storage infrastructure, rack integration, fleet health systems, telemetry platforms

Responsibilities

Drive technical triage and resolution of complex hardware failures; Partner with Fleet Health Engineering to investigate recurring issues and improve reliability; Lead root cause analysis for critical incidents; Coordinate repairs and lifecycle activities with vendors; Establish hardware maintenance standards and operational runbooks; Analyze failure trends to identify risks; Support new hardware introductions and validation; Coordinate spare parts strategy; Provide field feedback to improve platform designs; Mentor on-site technicians

Seniority

Senior, hands-on IC with leadership responsibilities

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗