CareerPlanSign in

(GPU) Infrastructure Engineer

🌐 Remote💼 Full-time🗓 2026-07-16 → 2026-09-26

Core

Design, build, and operate bare-metal GPU server fleets and the underlying infrastructure for on-prem and air-gapped data intelligence platforms.

Role type

Senior Infrastructure Engineer (GPU/HPC)

Builds

Production-ready GPU server environments and Kubernetes substrates for inference workloads

Domain

Data center infrastructure, HPC, GPU computing, air-gapped deployments

Deliverable

infrastructure

Required skills

Bare-metal Linux administration, NVIDIA GPU stack management, Zero-touch provisioning, Data center networking, Storage systems, On-site deployment execution

Preferred skills

NVIDIA DGX/HGX experience, InfiniBand/RDMA fabrics, Inference optimization, Field engineering

Technologies

Kubernetes, NVIDIA CUDA/GPU Operator/MIG/DCGM, Ansible, Terraform, Pulumi, Ceph, ZFS, NVMe, RDMA, Triton, KServe, vLLM

Responsibilities

Provision and operate bare-metal GPU server fleets with zero-touch automation; Tune NVIDIA GPU stacks for inference performance; Engineer resilient data-center networking and storage; Execute on-site server rack integration and commissioning; Partner with ML teams on on-prem inference serving; Plan capacity and run operational handovers

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.