CareerPlanSign in

Senior Debug System Engineer, Datacenter

US, CA, Santa Clara💼 Full-time💰 $168,000–$168,000🗓 2026-05-26 → 2026-10-07

Core

Drive failure analysis and debug efforts for GPU datacenter products (DGX, MGX, HGX) during the New Product Introduction phase to ensure quality transfer to mass production.

Role type

Senior System Debug Engineer (Hardware/Server)

Builds

GPU datacenter server products (DGX, MGX, HGX)

Domain

Semiconductor / Datacenter Hardware

Required skills

Failure analysis, root cause analysis, hardware debugging, server debugging, board-level debugging, log analysis, DFx enabling, experiment design, data analysis, report writing, vendor negotiation, time management

Preferred skills

Knowledge of sophisticated characterization equipment (oscilloscopes, analyzers), understanding of Hardware/Software/Firmware interactions, problem-solving mentality (via careerplan.io/jobs/893395268108-senior-debug-system-engineer-datacenter-at-nvidia)

Technologies

GPU baseboards, servers, rack-level systems, oscilloscopes, analyzers

Responsibilities

Perform failure analysis on GPU baseboards and servers at rack, system, and component levels; Analyze logs and failures spanning HW/SW/FW to propose debug strategies; Build experiments and collect/analyze data for root cause; Engage in DFx enabling efforts; Develop debug guides for partner teams; Communicate with vendors, suppliers, and internal/external teams

Seniority

Senior, hands-on IC