Principal Firmware Engineer – Server Manageability and Observability
Core
Lead end-to-end system software architecture for NVIDIA data center platforms (DGX, HGX), focusing on server manageability, observability, and SW/HW interfaces for GPUs, DPUs, and FPGAs.
Role type
Principal System Software Architect (Firmware & Server Management)
Builds
Next-generation data center products and system management protocols for hyperscalers and cloud providers.
Domain
Data Center Infrastructure, High-Performance Computing (HPC), Server Firmware & Management
Deliverable
production ML models | product features | infrastructure
Required skills
Scalable server system architecture, System firmware (SBIOS, OpenBMC), Linux kernel internals, Out-of-Band/In-Band management architectures, Device management protocols (MCTP, PLDM, SPDM, RDE), System management protocols (Redfish, IPMI), Networking technologies (TCP/IP, Ethernet, InfiniBand), Cross-functional project leadership, Left-shift strategy implementation
Preferred skills
Cloud and cluster level deployment management, Standards body participation (OCP, DMTF), NVIDIA HPC programming models (CUDA, cuDNN, DOCA), Enterprise storage architectures, Distributed parallel processing paradigms
Technologies
SBIOS, OpenBMC, Linux Kernel, MCTP, PLDM, SPDM, RDE, Redfish, IPMI, TCP/IP, Ethernet, InfiniBand, CUDA, cuDNN, DOCA
Responsibilities
Serve as primary technical point of contact for major customers and hyperscalers; Lead technical innovation and strategic collaborations for next-gen data center products; Align NVIDIA roadmap with customer requirements; Develop and drive adoption of new technologies and protocols; Make critical technical decisions in ambiguous situations using left-shift strategies.
Seniority
Principal, strategy & mentorship
