Staff Software Engineer, Inference / Compute Infrastructure Engineering
Core
Build software state machines that provision, manage, and decommission GPU hardware to create self-service inference clusters for the Research and Inference team.
Role type
Staff Software Engineer, Inference/Compute Infrastructure
Builds
Production-grade infrastructure-as-software platforms, declarative APIs, and control planes for GPU cluster lifecycle management.
Domain
AI Infrastructure, GPU Compute, Cloud Native Systems
Deliverable
production ML models | infrastructure
Required skills
Go, Python, or Rust; Durable workflow orchestration (Temporal, Cadence); State modeling and reconciliation (Kubernetes controllers); Event-driven systems (Kafka, NATS, SQS); Product mindset for internal platforms
Preferred skills
Bare-metal provisioning (PXE, Redfish, BMC); GPU stack knowledge (NCCL, CUDA, InfiniBand); Hyperscaler or GPU cloud experience; Systems programming in Rust or Go
Technologies
Go, Python, Rust, Temporal, Cadence, Kubernetes, Kafka, NATS, SQS, NCCL, CUDA, InfiniBand, RoCE
Responsibilities
Design and implement provisioning state machines for full hardware lifecycle; Build self-service APIs for cluster scaling and teardown; Automate self-healing and node replacement; Ensure pipeline reliability via idempotency and drift detection; Partner with ML teams to encode cluster topology constraints; Engineer infrastructure as typed, tested software with CI/CD
Seniority
Staff, hands-on IC with strategic ownership