Senior Software Engineer, Compute Architecture
Core
Build the software control plane for hardware lifecycle management across large-scale GPU data centers.
Role type
Senior IC backend engineer (infrastructure)
Builds
Go-based distributed services for hardware discovery, health monitoring, and operational workflows
Domain
Data center infrastructure / GPU compute / Distributed systems
Deliverable
production ML models | infrastructure
Required skills
Go, gRPC, REST APIs, Kubernetes, observability tooling (Prometheus, Grafana)
Preferred skills
GPU-based systems, low-level hardware management (BMC, Redfish), large-scale distributed systems, open-source contributions
Technologies
Go, gRPC, REST, Kubernetes, Prometheus, Grafana, BMC, Redfish
Responsibilities
Design and operate Go-based services managing GPU data center infrastructure lifecycle; Build automation for data center bring-up, hardware discovery, health monitoring, and remediation; Develop reliable APIs and workflows for managing BMCs, firmware state, and rack-level infrastructure; Improve observability and alerting tooling for production issues; Translate hardware failure modes into software improvements for platform resilience; Partner with hardware and operations teams to design fleet-scale systems
Seniority
Senior, hands-on IC