System Software Engineer, Node & Cluster Management
Core
Design and build node-level and cluster management software for AI silicon systems, including management planes, failover algorithms, and low-level utilities.
Role type
Senior System Software Engineer (Node & Cluster Management)
Builds
Node management APIs, cluster management solutions, CLI utilities, and firmware orchestration tools for AI hardware.
Domain
AI Infrastructure / High-Performance Computing / Silicon Management
Deliverable
production ML models | product features | infrastructure
Required skills
Linux systems development, C programming, Go/Rust/C++/Python, HTTP/REST API design, CLI tooling, kernel driver debugging, firmware integration, hardware bring-up, telemetry systems
Preferred skills
Redfish, OpenBMC, gNMI, IPMI, GPU/accelerator fleet management, lab automation, secure boot, attestation flows
Technologies
Linux, C, Go, Rust, Python, HTTP/REST, Redfish, OpenBMC, BMC firmware
Responsibilities
Design node-level management plane exposing health/telemetry via HTTP/REST; Implement cluster management solutions and failover algorithms; Build CLI utilities for operators to manage nodes and daemons; Partner with firmware engineers for unified in-band/out-of-band management; Extend capabilities to fleet-wide health aggregation and alerting; Debug production issues across APIs, daemons, drivers, and hardware; Build automation for lab system provisioning and test orchestration.
Seniority
Senior, hands-on IC