Software Engineer, GPU Infrastructure - HPC
Core
Build and maintain automation systems for provisioning, monitoring, and managing large-scale GPU server fleets to ensure high availability and performance for AI research and product development.
Role type
Senior IC infrastructure engineer (GPU/HPC systems)
Builds
Automated provisioning and management systems for server fleets; monitoring tools for health and lifecycle events
Domain
High-performance computing (HPC), distributed systems, data center infrastructure
Deliverable
infrastructure
Required skills
Python, Go, Linux, networking, server hardware management, SQL, PromQL, Pandas
Preferred skills
Low-level hardware component knowledge (PCIe, Infiniband), hardware management protocols (IPMI, Redfish), HPC or distributed systems experience, hardware design/management experience, monitoring tools (Prometheus, Grafana)
Responsibilities
Build automation for server fleet provisioning and management; Develop tools to monitor server health and performance; Identify and fix performance bottlenecks; Collaborate with clusters, networking, and infrastructure teams; Partner with external operators; Continuously improve automation to reduce manual work
Seniority
Senior, hands-on IC