Senior Staff Data Center Operations Engineer, GPU Hardware Architecture
Core
Senior Staff Data Center Operations Engineer acting as the definitive technical authority on GPU platforms, bridging silicon roadmaps with data center design and enabling predictive operations for high-density AI clusters.
Role type
Senior Staff IC Data Center Operations Engineer (GPU Hardware Architecture)
Builds
Next-gen liquid-cooled data center facilities, predictive maintenance tooling, and operational SOPs for GPU clusters.
Domain
AI Infrastructure / Data Center Operations / GPU Hardware Architecture
Deliverable
production ML models | product features | infrastructure
Required skills
NVIDIA/AMD GPU architecture mastery, NVLink/NVSwitch/InfiniBand expertise, Python/Go/Bash for telemetry tools, thermal management (D2C cooling), Root Cause Analysis (RCA), failure telemetry analysis, cross-functional consulting
Preferred skills
Experience with HBM4/Blackwell/Rubin roadmaps, liquid-cooling fluid dynamics, vendor technical audit
Technologies
NVIDIA Hopper/Blackwell/Rubin, AMD Instinct, NVLink, NVSwitch, InfiniBand, DCGM, ROCm, Python, Go, Bash
Responsibilities
Translate GPU power/thermal roadmaps into facility design requirements; develop diagnostic tools and SOPs for field technicians; architect site-level sparing strategies using failure telemetry; lead Tier-3 escalation and RCA for complex hardware failures; maintain 24-month silicon roadmap visibility; audit OEM/VAR hardware builds.
Seniority
Senior Staff, hands-on IC with strategic influence