Director of AI Infrastructure
Core
Oversee the full lifecycle of high-performance computing (HPC) environments, including on-prem GPU clusters and hybrid cloud orchestration, to power frontier AI research.
Role type
Director of AI Infrastructure
Builds
High-performance GPU clusters, hybrid cloud orchestration platforms (Beaker), and distributed storage systems for large-scale model training.
Domain
AI Research Infrastructure / High-Performance Computing
Deliverable
infrastructure
Required skills
Linux kernel, container runtimes, distributed systems, InfiniBand topologies, NCCL optimizations, Kubernetes, Slurm, distributed filesystems (WEKA, Ceph, Lustre), Go, Python
Preferred skills
None stated
Technologies
NVIDIA GPUs, InfiniBand, RoCE, AWS, GCP, Beaker, WEKA, Ceph, Lustre
Responsibilities
Oversee availability and performance of dense on-prem GPU clusters; Direct strategy for internal orchestration platform Beaker; Develop long-term roadmap for storage architecture; Act as primary steward of GPU compute budget; Serve as technical bridge to research teams.
Seniority
Director, leadership of multi-disciplinary engineering teams