CareerPlanGet AI match score →

Senior HPC Engineer, GPU Compute

💼 Full-time🗓 2026-07-14 → 2026-07-31

Core

Optimizing GPU clusters, InfiniBand networks, and KVM/QEMU virtualization stacks for high-performance AI cloud infrastructure.

Role type

Senior HPC Cluster Engineer (GPU Compute)

Builds

Hyperscaler AI cloud platform supporting data and model training to production deployment

Domain

Cloud Infrastructure / High-Performance Computing / GPU Systems

Deliverable

production ML models | infrastructure

Required skills

System-level software development, Linux administration, Server architecture (PCIe, NICs, Kernel), Performance-oriented programming (C/C++, Go, Python)

Preferred skills

GPU end-to-end testing in cluster environments, HPC workload optimization, RDMA/RoCE/InfiniBand protocols, Software-Defined Networking, QEMU/KVM virtualization, Deep learning frameworks (PyTorch, TensorFlow), Collective communication libraries (MPI, NCCL)

Technologies

Kubernetes, QEMU, KVM, InfiniBand, MPI, NCCL, PyTorch, TensorFlow

Responsibilities

Tune performance of GPU clusters and InfiniBand networks, Analyze and troubleshoot root causes of GPU/InfiniBand issues, Integrate new GPU hardware into infrastructure, Enhance automation systems for proactive monitoring, Configure and manage GPU devices and InfiniBand fabrics

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗