CareerPlanGet AI match score →

Principal Platform Software Engineer - RAS

2 Locations💼 Full-time💰 $272,000–$272,000🗓 2026-07-12 → 2026-07-31

Core

Design rack-level solutions for next-generation scaling AI supercomputing platforms using NVIDIA GH200 superchips, focusing on fleet management, health monitoring, and fault remediation at scale.

Role type

Principal Platform Software Engineer (AI Infrastructure)

Builds

Next-generation scaling AI supercomputing platforms and fleet management solutions for HPC and generative AI workloads.

Domain

AI Computing / High-Performance Computing (HPC) / Supercomputing

Deliverable

production ML models | product features | infrastructure

Required skills

C/C++, Python, time series databases (InfluxDB, Prometheus), REST API design (Redfish), firmware architecture, system resource analysis, scalability engineering, server platform debugging, SCM (Git, Perforce), Jira

Preferred skills

Telemetry collection & analysis engines, notification systems (PagerDuty), Open Compute Project (OCP) / DMTF contributions, x86 or ARM system architecture, Confidential Compute, ML and multi-variable optimization techniques

Technologies

NVIDIA GH200, Grace, InfluxDB, Prometheus, Grafana, Redfish, Git, Perforce, Jira, PagerDuty

Responsibilities

Drive fleet management solutions for scaling AI infrastructure; define architecture for health monitoring and fault remediation; write architecture specs and design documents; conduct POCs to validate architecture; ensure product testing and QA integration; manage product lifecycle and act as product owner; articulate requirements and execute end-to-end development plans.

Seniority

Principal, hands-on IC with strategy & mentorship

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗