CareerPlanGet AI match score →

Engineering Manager, GPU Infrastructure

United States🌐 Remote💼 Full-time🗓 2026-07-21 → 2026-07-31

Core

Lead a team of engineers building and operating GPU superclusters to power frontier AI models and distributed training workloads.

Role type

Engineering Manager, GPU Infrastructure

Builds

GPU clusters, distributed training environments, and high-performance computing infrastructure for AI research

Domain

Artificial Intelligence / High-Performance Computing / Cloud Infrastructure

Deliverable

production ML models | infrastructure

Required skills

Team leadership and mentorship, Kubernetes at scale, GPU/TPU cluster management, Infrastructure-as-Code (Terraform, ArgoCD), Distributed systems, Cost optimization, Vendor management

Preferred skills

Experience with JAX, PyTorch, or TensorFlow, Prometheus/Grafana monitoring, Multi-cloud environments

Technologies

Kubernetes, Terraform, ArgoCD, Prometheus, Grafana, JAX, PyTorch, TensorFlow

Responsibilities

Lead and mentor a team of GPU infrastructure engineers, Define and execute technical roadmap for cluster deployment and scaling, Collaborate with AI researchers to translate infrastructure needs into solutions, Establish observability frameworks for GPU utilization and reliability, Drive cost optimization initiatives for GPU infrastructure, Manage vendor relationships and hardware/cloud service contracts

Seniority

Manager, hands-on technical leadership

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗