Director of AI Infrastructure Seattle, WA View job
Core
Oversee the full lifecycle of high-performance computing (HPC) environments, including on-prem GPU clusters and hybrid cloud orchestration, to power frontier AI research.
Role type
Director of AI Infrastructure
Builds
High-performance GPU clusters, orchestration platforms (Beaker), and distributed storage systems for AI research teams
Domain
AI Research Infrastructure / High-Performance Computing (HPC)
Deliverable
infrastructure
Required skills
Linux kernel management, distributed systems, InfiniBand topology optimization, Kubernetes, Slurm, Go, Python, NVIDIA GPU cluster management, WEKA, Ceph, Lustre
Preferred skills
Strategic system design for 3-5 year horizons, resource economics, hybrid-cloud architecture
Responsibilities
Manage availability and performance of dense on-prem GPU clusters; Direct strategy for internal orchestration platform (Beaker); Develop long-term storage architecture roadmap; Steward GPU compute budget and make cloud vs on-prem capacity decisions; Serve as technical bridge to research teams to ensure infrastructure velocity
Seniority
Director, leadership of multi-disciplinary engineering teams