CareerPlanSign in

Software Engineer, GPU Infrastructure - HPC

San Francisco💼 Full-time🗓 2026-02-05 → 2026-09-26

Core

Build and maintain automation systems for provisioning, monitoring, and managing large-scale GPU server fleets to ensure high availability and performance for AI research and product development.

Role type

Senior IC infrastructure engineer (GPU/HPC systems)

Builds

Automated provisioning and management systems for server fleets; monitoring tools for health and lifecycle events

Domain

High-performance computing (HPC), distributed systems, data center infrastructure

Deliverable

infrastructure

Required skills

Python, Go, Linux, networking, server hardware management, SQL, PromQL, Pandas

Preferred skills

Low-level hardware component knowledge (PCIe, Infiniband), hardware management protocols (IPMI, Redfish), HPC or distributed systems experience, hardware design/management experience, monitoring tools (Prometheus, Grafana)

Responsibilities

Build automation for server fleet provisioning and management; Develop tools to monitor server health and performance; Identify and fix performance bottlenecks; Collaborate with clusters, networking, and infrastructure teams; Partner with external operators; Continuously improve automation to reduce manual work

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.