Senior HPC & GPU Infrastructure Engineer
Core
Own the health, reliability, and performance of a high-density GPU compute cluster for frontier AI models and real-time applications. (via careerplan.io/jobs/840adeda-38a1-4ca4-b7c9-304f651e2e55-senior-hpc-gpu-infrastructure-engineer-at-sciforium)
Role type
Senior HPC & GPU Infrastructure Engineer (SRE)
Builds
High-performance GPU clusters for multimodal AI models and inference workloads
Domain
AI Infrastructure / High-Performance Computing / GPU Systems
Deliverable
infrastructure
Required skills
Linux systems engineering, GPU driver bring-up, cluster monitoring, network security, distributed file systems, Bash scripting, Python automation, NVIDIA/AMD GPU debugging, kernel module management
Preferred skills
Job schedulers (Slurm, Kubernetes, Run:AI), vLLM, model serving optimizations, configuration management (Ansible, SaltStack, Terraform)
Technologies
Ubuntu, CentOS, RHEL, CUDA, ROCm, PyTorch, JAX, vLLM, NCCL, cuDNN, NVLink, RDMA, NFS, GPFS, Lustre, FreeIPA, LDAP, SSH, iptables, Slurm, Kubernetes, Run:AI, Ansible, SaltStack, Terraform
Responsibilities
Respond to system outages and GPU failures; implement cluster monitoring for GPU health and topology; coordinate with vendors for hardware repairs; manage Linux OS patching and kernel tuning; configure network security and identity management; deploy and integrate new GPU nodes; debug complex GPU/ML stack interactions
Seniority
Senior, hands-on IC
