Compute Infrastructure Lead
Core
Own and scale the compute backbone for training, evaluation, and data-processing workloads to enable model deployment from research to production.
Role type
Senior IC Compute Infrastructure Lead (ML Training & HPC)
Builds
Multi-provider GPU fleet, cloud scheduler, distributed training/eval frameworks, virtualized dev environments, and observability tooling.
Domain
AI/ML Infrastructure, High-Performance Computing (HPC), Cloud Engineering
Deliverable
production ML models | infrastructure
Required skills
Large-scale GPU cluster operations, distributed training frameworks (PyTorch, NCCL), cloud scheduling (Kubernetes, Slurm, Ray), high-performance networking (InfiniBand, RoCE), Python systems engineering, cost optimization, capacity planning.
Preferred skills
Online/continuous RL, multi-cloud fabrics (SkyPilot), robotics/autonomous vehicle training stacks, open-source contributions.
Technologies
Ray, Kubernetes, Slurm, SkyPilot, Prometheus, Grafana, MLflow, InfiniBand, RoCE, PyTorch, NCCL
Responsibilities
Provision GPU capacity across cloud providers and HPC; design and operate the cloud scheduler with dynamic checkpointing; build distributed compute frameworks for heterogeneous hardware; orchestrate data-processing workloads; deliver virtualized GPU/CPU dev sessions; build unified observability; manage capacity, cost, and provider relationships.
Seniority
Senior, hands-on IC with lead potential