Four positions at Stanford Research Computing
Core
Managing and maintaining large-scale HPC storage environments (Oak, Fir, Elm) and GPU clusters (Marlowe) for university researchers.
Role type
Senior Infrastructure & Storage Systems Engineer
Builds
PB-scale Lustre file storage, object storage on tape, NVIDIA DGX H100 SuperPOD clusters, and Infiniband/Ethernet HPC fabrics.
Domain
High-Performance Computing (HPC), Data Infrastructure, Storage Systems
Deliverable
infrastructure
Required skills
Lustre, Infiniband, PB-scale storage management, NVIDIA DGX/H100, DDN Intelliflash, HPC cluster administration, hardware troubleshooting, system architecture
Preferred skills
Tape storage management, SuperPOD operations, user/PI interaction
Technologies
Lustre, Infiniband, NVIDIA DGX H100, DDN Intelliflash, Ethernet
Responsibilities
Leading storage teams and setting direction for large storage environments, maintaining and expanding 20+ PB Lustre storage, keeping GPU/AI environments up-to-date, troubleshooting hardware and infrastructure, interacting with users and PIs
Seniority
Senior, hands-on IC / Technical Manager