Staff Software Engineer, GPU Infrastructure Lifecycle Management
Core
Build software state machines and control planes to automate the full lifecycle of GPU hardware, provisioning inference clusters via API without manual intervention.
Role type
Staff Software Engineer, GPU Infrastructure Lifecycle Management
Builds
Manifest-driven software systems for GPU cluster provisioning, scaling, and decommissioning
Domain
AI Infrastructure / GPU Cloud / Datacenter Operations
Deliverable
production ML models | infrastructure
Required skills
Go, Python, Rust, durable workflow orchestration (Temporal/Cadence), control plane design, event-driven systems (Kafka/NATS/SQS), product mindset for internal APIs
Preferred skills
bare-metal provisioning (PXE/iPXE/Redfish), networking (VLANs/BGP), GPU stack knowledge (NCCL/CUDA/InfiniBand), hyperscaler experience, Rust/Go systems programming
Technologies
Go, Python, Rust, Temporal, Cadence, Kafka, NATS, SQS, Kubernetes, PXE, iPXE, Redfish, IPMI, CUDA, NCCL, InfiniBand, RoCE
Responsibilities
Design and implement provisioning state machines for GPU hosts; Build self-service declarative APIs for cluster management; Automate self-healing and node replacement; Ensure pipeline reliability (idempotency, retries, drift detection); Partner with ML teams to encode cluster topology constraints; Engineer infrastructure as typed, tested software with CI/CD
Seniority
Staff, hands-on IC with strategic ownership