CareerPlanSign in

AI Infrastructure Engineer

San Francisco💼 Full-time🗓 2026-08-06 → 2026-09-26

Core

Own the entire software stack of GPU clusters, from kernel tuning and drivers through schedulers and containers, to serve foundation model training and model serving teams.

Role type

Senior AI Infrastructure Engineer (GPU Cluster & HPC)

Builds

Production-ready GPU nodes, automated provisioning pipelines, and high-performance training/serving clusters.

Domain

AI Infrastructure, High-Performance Computing (HPC), GPU Clusters

Deliverable

infrastructure

Required skills

Linux internals (kernel modules, cgroups, NUMA), GPU driver stacks (NVIDIA CUDA, AMD ROCm), Kubernetes (GPU workloads), HPC schedulers (Slurm), Configuration Management (Ansible/SaltStack), Provisioning tooling (Packer/MaaS), Python/Bash automation, Distributed filesystems (Lustre/GPFS), NCCL/RDMA networking, Container runtime (Docker/containerd)

Preferred skills

Foundation model training support, Inference serving stacks (vLLM/Triton/TensorRT), GPU profiling tools (Nsight/eBPF)

Technologies

Kubernetes, Slurm, Ansible, SaltStack, Packer, Terraform, NVIDIA CUDA, AMD ROCm, PyTorch, JAX, InfiniBand, RoCE, Prometheus, Grafana, DCGM

Responsibilities

Author versioned OS images and automated pipelines for node bring-up; Build automated acceptance suites for node validation; Execute rolling upgrades and maintain driver/framework compatibility; Automate detection and remediation of unhealthy nodes; Manage cluster configuration via IaC and Git workflows; Operate GPU-enabled Kubernetes and training schedulers; Debug complex distributed system issues and maintain observability.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.