CareerPlanSign in

Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure

Poland, Warsaw💼 Full-time💰 $221,250–$221,250🗓 2026-07-22 → 2026-09-26

Core

Lead the design, deployment, and reliability of large-scale GPU compute clusters running demanding deep learning training, inference, and HPC workloads.

Role type

Senior HPC Cluster Administrator (Deep Learning Infrastructure)

Builds

Large-scale GPU compute clusters (DGX/HGX, Grace Blackwell) supporting distributed training and inference

Domain

High-Performance Computing / Deep Learning Infrastructure

Deliverable

production ML models | infrastructure

Required skills

Linux systems administration at scale, GPU cluster lifecycle management, storage solution design (NFS, Lustre, WekaFS), infrastructure automation (Ansible, Terraform), job scheduling (Slurm), observability (Prometheus, Grafana, DCGM), high-speed networking (InfiniBand, RDMA), container technologies (Docker, Apptainer, Kubernetes), Python/Bash scripting

Preferred skills

NVIDIA GPU infrastructure tools (DCGM, MIG), cluster management platforms (Colossus, Bright Cluster Manager), MLOps tooling, BMC/IPMI/Redfish management

Technologies

DGX, HGX, Grace Blackwell, InfiniBand, NVLink, EFA/RDMA, NFS, Lustre, WekaFS, Ansible, Terraform, GitLab, Slurm, Prometheus, Grafana, DCGM, Docker, Apptainer, Kubernetes, PyTorch, JAX, Megatron

Responsibilities

Own full lifecycle of GPU compute clusters (procurement to deprecation), design and scale storage solutions, lead automation using IaC and CI/CD, manage and optimize job scheduling via Slurm, maintain observability stacks and resolve incidents, collaborate with ML engineers to tune cluster configurations, evaluate and introduce new technologies, mentor junior engineers

Seniority

Senior, hands-on IC with mentorship responsibilities

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.