CareerPlanSign in

Senior HPC & GPU Infrastructure Engineer

San Francisco💼 Full-time🗓 2026-05-08 → 2026-10-07

Core

Own the health, reliability, and performance of a high-density GPU compute cluster for frontier AI models and real-time applications. (via careerplan.io/jobs/840adeda-38a1-4ca4-b7c9-304f651e2e55-senior-hpc-gpu-infrastructure-engineer-at-sciforium)

Role type

Senior HPC & GPU Infrastructure Engineer (SRE)

Builds

High-performance GPU clusters for multimodal AI models and inference workloads

Domain

AI Infrastructure / High-Performance Computing / GPU Systems

Deliverable

infrastructure

Required skills

Linux systems engineering, GPU driver bring-up, cluster monitoring, network security, distributed file systems, Bash scripting, Python automation, NVIDIA/AMD GPU debugging, kernel module management

Preferred skills

Job schedulers (Slurm, Kubernetes, Run:AI), vLLM, model serving optimizations, configuration management (Ansible, SaltStack, Terraform)

Technologies

Ubuntu, CentOS, RHEL, CUDA, ROCm, PyTorch, JAX, vLLM, NCCL, cuDNN, NVLink, RDMA, NFS, GPFS, Lustre, FreeIPA, LDAP, SSH, iptables, Slurm, Kubernetes, Run:AI, Ansible, SaltStack, Terraform

Responsibilities

Respond to system outages and GPU failures; implement cluster monitoring for GPU health and topology; coordinate with vendors for hardware repairs; manage Linux OS patching and kernel tuning; configure network security and identity management; deploy and integrate new GPU nodes; debug complex GPU/ML stack interactions

Seniority

Senior, hands-on IC