Manager, HPC Storage Engineer
Core
Lead the team responsible for Runpod's distributed storage infrastructure across all regions, owning the end-to-end storage stack from NAND/NVMe devices through filesystems and transport protocols to ensure performance and scalability for AI workloads.
Role type
Engineering Manager, Datacenter Storage Engineering
Builds
Global storage platforms supporting training, inference, checkpointing, and dataset access at scale
Domain
Cloud infrastructure for AI/ML, HPC storage systems
Deliverable
infrastructure
Required skills
Engineering leadership, distributed storage architecture, VAST Data deployment, parallel filesystems (Lustre/GPFS/BeeGFS), low-level storage stack knowledge (NAND/NVMe/PCIe), high-performance data paths (NFS over RDMA), Linux internals, operational excellence
Preferred skills
AI training pipeline support, RDMA fabrics, multi-tenant isolation design, hyperscale/HPC/AI infrastructure background
Technologies
VAST Data, Lustre, GPFS, BeeGFS, NFS, RDMA, NVMe, NAND
Responsibilities
Define and operate global storage platforms; manage and grow storage engineering teams; design and operate large-scale SAN and NFS deployments; lead deployments of VAST Data and parallel filesystems; drive performance optimization from media through networking; evaluate and deploy next-gen capabilities like GPU Direct Storage; establish best practices for reliability and scale; build automation for provisioning and monitoring; partner with cross-functional teams; manage vendor and partner relationships
Seniority
Manager, hands-on technical leadership