AI Infrastructure Engineer
Core
Own the entire software stack of GPU clusters, from kernel tuning and drivers through schedulers and containers, to serve foundation model training and model serving teams.
Role type
Senior AI Infrastructure Engineer (GPU Cluster & HPC)
Builds
Production-ready GPU nodes, automated provisioning pipelines, and high-performance training/serving clusters.
Domain
AI Infrastructure, High-Performance Computing (HPC), GPU Clusters
Deliverable
infrastructure
Required skills
Linux internals (kernel modules, cgroups, NUMA), GPU driver stacks (NVIDIA CUDA, AMD ROCm), Kubernetes (GPU workloads), HPC schedulers (Slurm), Configuration Management (Ansible/SaltStack), Provisioning tooling (Packer/MaaS), Python/Bash automation, Distributed filesystems (Lustre/GPFS), NCCL/RDMA networking, Container runtime (Docker/containerd)
Preferred skills
Foundation model training support, Inference serving stacks (vLLM/Triton/TensorRT), GPU profiling tools (Nsight/eBPF)
Technologies
Kubernetes, Slurm, Ansible, SaltStack, Packer, Terraform, NVIDIA CUDA, AMD ROCm, PyTorch, JAX, InfiniBand, RoCE, Prometheus, Grafana, DCGM
Responsibilities
Author versioned OS images and automated pipelines for node bring-up; Build automated acceptance suites for node validation; Execute rolling upgrades and maintain driver/framework compatibility; Automate detection and remediation of unhealthy nodes; Manage cluster configuration via IaC and Git workflows; Operate GPU-enabled Kubernetes and training schedulers; Debug complex distributed system issues and maintain observability.
Seniority
Senior, hands-on IC