Staff Software Engineer, DC Infrastructure
Core
Develop software for managing GPU server fleets and data centers, focusing on diagnostics, observability, automation, and repair tooling for high-performance compute clusters.
Role type
Staff Software Engineer (Data Center Infrastructure)
Builds
Automation tooling, AI agents for hardware diagnosis/remediation, monitoring systems, and facilities management software for power and liquid cooling.
Domain
AI Infrastructure / Data Center Operations / High-Performance Computing
Deliverable
production ML models | product features | infrastructure
Required skills
Distributed systems, reliability engineering, cloud platforms (Kubernetes, IaC, GCP), Go/Python/Java/Rust, hardware diagnostics, automation development, post-repair validation testing.
Preferred skills
Temporal, Kubernetes, hardware vendor collaboration, large-scale GPU fleet operations, hyperscale data center experience.
Technologies
Kubernetes, GCP, Temporal, NVIDIA A100/H200/GB200/B200, AMD 350X/355X, PyTorch, NVIDIA NCCL, direct liquid cooling systems.
Responsibilities
Develop deep-level diagnostics and troubleshooting for hardware faults in GPU racks; build automation tooling for GPU platforms; create AI agents for component-level diagnosis and remediation; develop tooling for critical environment management; build post-repair validation tools; own deployment and operational support of tooling; develop automation for facilities power and liquid cooling hardware.
Seniority
Staff, hands-on IC with technical direction