Principal Platform Software Engineer - RAS
Core
Design rack-level solutions for next-generation scaling AI supercomputing platforms using NVIDIA GH200 superchips, focusing on fleet management, health monitoring, and fault remediation at scale.
Role type
Principal Platform Software Engineer (AI Infrastructure)
Builds
Next-generation scaling AI supercomputing platforms and fleet management solutions for HPC and generative AI workloads.
Domain
AI Computing / High-Performance Computing (HPC) / Supercomputing
Deliverable
production ML models | product features | infrastructure
Required skills
C/C++, Python, time series databases (InfluxDB, Prometheus), REST API design (Redfish), firmware architecture, system resource analysis, scalability engineering, server platform debugging, SCM (Git, Perforce), Jira
Preferred skills
Telemetry collection & analysis engines, notification systems (PagerDuty), Open Compute Project (OCP) / DMTF contributions, x86 or ARM system architecture, Confidential Compute, ML and multi-variable optimization techniques
Technologies
NVIDIA GH200, Grace, InfluxDB, Prometheus, Grafana, Redfish, Git, Perforce, Jira, PagerDuty
Responsibilities
Drive fleet management solutions for scaling AI infrastructure; define architecture for health monitoring and fault remediation; write architecture specs and design documents; conduct POCs to validate architecture; ensure product testing and QA integration; manage product lifecycle and act as product owner; articulate requirements and execute end-to-end development plans.
Seniority
Principal, hands-on IC with strategy & mentorship