Engineering Manager, GPU Infrastructure
Core
Lead a team of engineers building and operating GPU superclusters to power frontier AI models and distributed training workloads.
Role type
Engineering Manager, GPU Infrastructure
Builds
GPU clusters, distributed training environments, and high-performance computing infrastructure for AI research
Domain
Artificial Intelligence / High-Performance Computing / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Team leadership and mentorship, Kubernetes at scale, GPU/TPU cluster management, Infrastructure-as-Code (Terraform, ArgoCD), Distributed systems, Cost optimization, Vendor management
Preferred skills
Experience with JAX, PyTorch, or TensorFlow, Prometheus/Grafana monitoring, Multi-cloud environments
Technologies
Kubernetes, Terraform, ArgoCD, Prometheus, Grafana, JAX, PyTorch, TensorFlow
Responsibilities
Lead and mentor a team of GPU infrastructure engineers, Define and execute technical roadmap for cluster deployment and scaling, Collaborate with AI researchers to translate infrastructure needs into solutions, Establish observability frameworks for GPU utilization and reliability, Drive cost optimization initiatives for GPU infrastructure, Manage vendor relationships and hardware/cloud service contracts
Seniority
Manager, hands-on technical leadership