CareerPlanGet AI match score →

Staff Software Engineer, GPU Infrastructure Lifecycle Management

San Francisco💼 Full-time💰 $240,000–$240,000🗓 2026-07-16 → 2026-07-31

Core

Build software state machines and control planes to automate the full lifecycle of GPU hardware, provisioning inference clusters via API without manual intervention.

Role type

Staff Software Engineer, GPU Infrastructure Lifecycle Management

Builds

Manifest-driven software systems for GPU cluster provisioning, scaling, and decommissioning

Domain

AI Infrastructure / GPU Cloud / Datacenter Operations

Deliverable

production ML models | infrastructure

Required skills

Go, Python, Rust, durable workflow orchestration (Temporal/Cadence), control plane design, event-driven systems (Kafka/NATS/SQS), product mindset for internal APIs

Preferred skills

bare-metal provisioning (PXE/iPXE/Redfish), networking (VLANs/BGP), GPU stack knowledge (NCCL/CUDA/InfiniBand), hyperscaler experience, Rust/Go systems programming

Technologies

Go, Python, Rust, Temporal, Cadence, Kafka, NATS, SQS, Kubernetes, PXE, iPXE, Redfish, IPMI, CUDA, NCCL, InfiniBand, RoCE

Responsibilities

Design and implement provisioning state machines for GPU hosts; Build self-service declarative APIs for cluster management; Automate self-healing and node replacement; Ensure pipeline reliability (idempotency, retries, drift detection); Partner with ML teams to encode cluster topology constraints; Engineer infrastructure as typed, tested software with CI/CD

Seniority

Staff, hands-on IC with strategic ownership

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗