Director of Infrastructure Engineering
Core
Lead and scale Runpod's core cloud and bare-metal environments, owning critical foundational layers including SRE, global networking, HPC networks, and distributed storage engines to support massive GPU computing demands.
Role type
Director of Infrastructure Engineering (SRE, Networking, Storage)
Builds
High-availability cloud and bare-metal infrastructure for AI/ML workloads
Domain
AI Infrastructure / Cloud Computing / HPC
Deliverable
production ML models | infrastructure
Required skills
Engineering leadership (managing managers), distributed systems architecture, HPC networking (InfiniBand, RoCE), distributed storage systems, SRE practices, infrastructure as code, container orchestration
Preferred skills
GPU cluster architecture, hardware/GPU interconnects knowledge, hyper-growth startup scaling, open-source contributions
Technologies
InfiniBand, RoCE, Ceph, Lustre, Weka, NVMe-oF, Terraform, Ansible, Kubernetes, BGP
Responsibilities
Lead SRE, networking, and storage teams; Architect global network backbone and HPC cluster networks; Direct storage engine architecture and performance tuning; Build and mentor high-output engineering orgs; Translate scale challenges into technical roadmaps; Improve infrastructure reliability and delivery metrics; Provide architectural oversight for bare-metal and virtualization layers
Seniority
Director, hands-on IC with management responsibilities