(GPU) Infrastructure Engineer
Core
Design, build, and operate bare-metal GPU server fleets and the underlying infrastructure for on-prem and air-gapped data intelligence platforms.
Role type
Senior Infrastructure Engineer (GPU/HPC)
Builds
Production-ready GPU server environments and Kubernetes substrates for inference workloads
Domain
Data center infrastructure, HPC, GPU computing, air-gapped deployments
Deliverable
infrastructure
Required skills
Bare-metal Linux administration, NVIDIA GPU stack management, Zero-touch provisioning, Data center networking, Storage systems, On-site deployment execution
Preferred skills
NVIDIA DGX/HGX experience, InfiniBand/RDMA fabrics, Inference optimization, Field engineering
Technologies
Kubernetes, NVIDIA CUDA/GPU Operator/MIG/DCGM, Ansible, Terraform, Pulumi, Ceph, ZFS, NVMe, RDMA, Triton, KServe, vLLM
Responsibilities
Provision and operate bare-metal GPU server fleets with zero-touch automation; Tune NVIDIA GPU stacks for inference performance; Engineer resilient data-center networking and storage; Execute on-site server rack integration and commissioning; Partner with ML teams on on-prem inference serving; Plan capacity and run operational handovers
Seniority
Senior, hands-on IC