Head of Engineering - GPU Cloud
Core
Lead Support Engineering, HPC, and SRE teams to design, deploy, and operate large-scale GPU clusters for AI and HPC workloads.
Role type
Senior Engineering Leadership (Head of GPU Cloud)
Builds
Large-scale sovereign GPU cloud infrastructure for AI and HPC
Domain
Cloud Infrastructure / High-Performance Computing / AI
Deliverable
production ML models | infrastructure
Required skills
Infrastructure engineering leadership, large-scale cluster design and operations, GPU/HPC environment expertise, distributed infrastructure architecture, orchestration and provisioning, observability, high-performance storage, architectural trade-off analysis
Preferred skills
NVIDIA and AMD technologies, Kubernetes, Proxmox, Warewulf, Prometheus, Grafana, Lustre, DDN, VAST Data
Technologies
Kubernetes, Proxmox, Warewulf, Prometheus, Grafana, Lustre, DDN, VAST Data, NVIDIA, AMD
Responsibilities
Lead Support Engineering, HPC, and SRE organizations; Own technical strategy and architecture for GPU Cloud; Validate technical feasibility of commercial proposals; Oversee design and deployment of new GPU clusters; Drive evolution of cluster and capacity management; Ensure reliability and scalability of GPU infrastructure; Provide technical leadership on complex AI/HPC projects; Build alignment between Engineering, GTM, Product, and Operations; Develop the engineering organization.
Seniority
Senior, hands-on IC with management responsibilities