CareerPlanSign in

Senior Staff Data Center Operations Engineer, GPU Hardware Architecture

San Francisco, CA - US💼 Full-time💰 $179,000–$179,000🗓 2026-05-15 → 2026-09-26

Core

Senior Staff Data Center Operations Engineer acting as the definitive technical authority on GPU platforms, bridging silicon roadmaps with data center design and enabling predictive operations for high-density AI clusters.

Role type

Senior Staff IC Data Center Operations Engineer (GPU Hardware Architecture)

Builds

Next-gen liquid-cooled data center facilities, predictive maintenance tooling, and operational SOPs for GPU clusters.

Domain

AI Infrastructure / Data Center Operations / GPU Hardware Architecture

Deliverable

production ML models | product features | infrastructure

Required skills

NVIDIA/AMD GPU architecture mastery, NVLink/NVSwitch/InfiniBand expertise, Python/Go/Bash for telemetry tools, thermal management (D2C cooling), Root Cause Analysis (RCA), failure telemetry analysis, cross-functional consulting

Preferred skills

Experience with HBM4/Blackwell/Rubin roadmaps, liquid-cooling fluid dynamics, vendor technical audit

Technologies

NVIDIA Hopper/Blackwell/Rubin, AMD Instinct, NVLink, NVSwitch, InfiniBand, DCGM, ROCm, Python, Go, Bash

Responsibilities

Translate GPU power/thermal roadmaps into facility design requirements; develop diagnostic tools and SOPs for field technicians; architect site-level sparing strategies using failure telemetry; lead Tier-3 escalation and RCA for complex hardware failures; maintain 24-month silicon roadmap visibility; audit OEM/VAR hardware builds.

Seniority

Senior Staff, hands-on IC with strategic influence

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.