CareerPlanSign in

Staff Software Engineer, DC Infrastructure

San Francisco, CA - US💼 Full-time💰 $215,000–$215,000🗓 2026-08-12 → 2026-09-25

Core

Develop software for managing GPU server fleets and data centers, focusing on diagnostics, observability, automation, and repair tooling for high-performance compute clusters.

Role type

Staff Software Engineer (Data Center Infrastructure)

Builds

Automation tooling, AI agents for hardware diagnosis/remediation, monitoring systems, and facilities management software for power and liquid cooling.

Domain

AI Infrastructure / Data Center Operations / High-Performance Computing

Deliverable

production ML models | product features | infrastructure

Required skills

Distributed systems, reliability engineering, cloud platforms (Kubernetes, IaC, GCP), Go/Python/Java/Rust, hardware diagnostics, automation development, post-repair validation testing.

Preferred skills

Temporal, Kubernetes, hardware vendor collaboration, large-scale GPU fleet operations, hyperscale data center experience.

Technologies

Kubernetes, GCP, Temporal, NVIDIA A100/H200/GB200/B200, AMD 350X/355X, PyTorch, NVIDIA NCCL, direct liquid cooling systems.

Responsibilities

Develop deep-level diagnostics and troubleshooting for hardware faults in GPU racks; build automation tooling for GPU platforms; create AI agents for component-level diagnosis and remediation; develop tooling for critical environment management; build post-repair validation tools; own deployment and operational support of tooling; develop automation for facilities power and liquid cooling hardware.

Seniority

Staff, hands-on IC with technical direction

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.