Senior Engineering Manager, AI Infrastructure
Core
Run the day-to-day operation of high-performance computing (HPC) systems, including on-prem GPU clusters and hybrid cloud orchestration, to power AI research.
Role type
Senior Engineering Manager, AI Infrastructure
Builds
Reliable, high-utilization HPC platforms for frontier AI model training and research
Domain
AI Research Infrastructure / High-Performance Computing
Deliverable
infrastructure
Required skills
Linux kernel, container runtimes, distributed systems, InfiniBand topologies, NCCL optimizations, Kubernetes, Slurm, distributed filesystems, Go, Python
Preferred skills
None stated
Technologies
NVIDIA GPUs, InfiniBand, RoCE, AWS, GCP, Beaker, WEKA, Ceph, Lustre
Responsibilities
Manage availability and performance of dense on-prem GPU clusters; Operate and improve internal orchestration platform (Beaker) for resource allocation; Execute and improve storage environment for petascale data; Manage GPU compute allocation against budget; Serve as technical bridge to research teams; Manage and grow a team of systems engineers and developers
Seniority
Senior, hands-on IC with team leadership