Senior Infrastructure Engineer - GPU Compute
Core
Build and operate a large, heterogeneous, globally distributed GPU fleet to power AI inference workloads, ensuring high utilization, low cost, and reliability.
Role type
Senior Infrastructure Engineer (GPU Compute)
Builds
A global compute fabric for AI inference using consumer and datacenter GPUs
Domain
AI Infrastructure / GPU Compute / Cloud & On-Prem Hybrid
Deliverable
production ML models
Required skills
Kubernetes, Docker, Linux systems administration, Infrastructure-as-Code (Terraform, Ansible, Pulumi), Python/Bash/TypeScript/Go, GPU fleet orchestration, spot instance management, observability, secure fleet access
Preferred skills
CUDA, bare-metal optimization, kernel tuning, Ray/Slurm, network topology design, Rust, on-prem data center operations
Technologies
SkyPilot, Kubernetes, k3s, Tailscale, Teleport, RTX 5090, EPYC SP5
Responsibilities
Operate a heterogeneous multi-region GPU fleet; Maximize GPU utilization across inference workloads; Optimize GPU performance via PCIe P2P, ReBAR, NUMA, and driver tuning; Build secure fleet access and robust observability; Drive down cost per GPU-hour through spot management and intelligent placement
Seniority
Senior, hands-on IC