Senior ML Infrastructure Engineer
Core
Own and evolve multi-cluster GPU infrastructure for training tabular foundation models, optimizing cost, reliability, and throughput for world-leading AI research.
Role type
Senior IC ML Infrastructure Engineer (GPU/Cluster)
Builds
Multi-cluster GPU orchestration, distributed training systems, and developer productivity tooling for tabular foundation models.
Domain
AI/ML Infrastructure, High-Performance Computing, Cloud Systems
Deliverable
infrastructure
Required skills
GPU infrastructure operations, distributed training systems, cluster management, systems-level debugging, Python, PyTorch internals, cost optimization, multi-cluster orchestration
Preferred skills
Multi-cloud/HPC experience, Triton/CUDA/custom kernels, experiment tracking tooling, scaling from single to multi-cluster
Technologies
Slurm, GCP, Docker, wandb, GitHub Actions, uv, PyTorch, Triton
Responsibilities
Own multi-cluster GPU infrastructure architecture and scheduling; Drive GPU utilization and training throughput via profiling and optimization; Architect next-gen infrastructure for new hardware and providers; Build developer productivity layer (CI, experiment tracking, model registry); Own compute budget and cost per FLOP optimization.
Seniority
Senior, hands-on IC