Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure
Core
Lead the design, deployment, and reliability of large-scale GPU compute clusters running demanding deep learning training, inference, and HPC workloads.
Role type
Senior HPC Cluster Administrator (Deep Learning Infrastructure)
Builds
Large-scale GPU compute clusters (DGX/HGX, Grace Blackwell) supporting distributed training and inference
Domain
High-Performance Computing / Deep Learning Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Linux systems administration at scale, GPU cluster lifecycle management, storage solution design (NFS, Lustre, WekaFS), infrastructure automation (Ansible, Terraform), job scheduling (Slurm), observability (Prometheus, Grafana, DCGM), high-speed networking (InfiniBand, RDMA), container technologies (Docker, Apptainer, Kubernetes), Python/Bash scripting
Preferred skills
NVIDIA GPU infrastructure tools (DCGM, MIG), cluster management platforms (Colossus, Bright Cluster Manager), MLOps tooling, BMC/IPMI/Redfish management
Technologies
DGX, HGX, Grace Blackwell, InfiniBand, NVLink, EFA/RDMA, NFS, Lustre, WekaFS, Ansible, Terraform, GitLab, Slurm, Prometheus, Grafana, DCGM, Docker, Apptainer, Kubernetes, PyTorch, JAX, Megatron
Responsibilities
Own full lifecycle of GPU compute clusters (procurement to deprecation), design and scale storage solutions, lead automation using IaC and CI/CD, manage and optimize job scheduling via Slurm, maintain observability stacks and resolve incidents, collaborate with ML engineers to tune cluster configurations, evaluate and introduce new technologies, mentor junior engineers
Seniority
Senior, hands-on IC with mentorship responsibilities