CareerPlanSign in

Senior ML Infrastructure Engineer

Berlin💼 Full-time🗓 2026-06-25 → 2026-09-26

Core

Own and evolve multi-cluster GPU infrastructure for training tabular foundation models, optimizing cost, reliability, and throughput for world-leading AI research.

Role type

Senior IC ML Infrastructure Engineer (GPU/Cluster)

Builds

Multi-cluster GPU orchestration, distributed training systems, and developer productivity tooling for tabular foundation models.

Domain

AI/ML Infrastructure, High-Performance Computing, Cloud Systems

Deliverable

infrastructure

Required skills

GPU infrastructure operations, distributed training systems, cluster management, systems-level debugging, Python, PyTorch internals, cost optimization, multi-cluster orchestration

Preferred skills

Multi-cloud/HPC experience, Triton/CUDA/custom kernels, experiment tracking tooling, scaling from single to multi-cluster

Technologies

Slurm, GCP, Docker, wandb, GitHub Actions, uv, PyTorch, Triton

Responsibilities

Own multi-cluster GPU infrastructure architecture and scheduling; Drive GPU utilization and training throughput via profiling and optimization; Architect next-gen infrastructure for new hardware and providers; Build developer productivity layer (CI, experiment tracking, model registry); Own compute budget and cost per FLOP optimization.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.