Senior HPC Engineer, GPU Compute
Core
Optimizing GPU clusters, InfiniBand networks, and KVM/QEMU virtualization stacks for high-performance AI cloud infrastructure.
Role type
Senior HPC Cluster Engineer (GPU Compute)
Builds
Hyperscaler AI cloud platform supporting data and model training to production deployment
Domain
Cloud Infrastructure / High-Performance Computing / GPU Systems
Deliverable
production ML models | infrastructure
Required skills
System-level software development, Linux administration, Server architecture (PCIe, NICs, Kernel), Performance-oriented programming (C/C++, Go, Python)
Preferred skills
GPU end-to-end testing in cluster environments, HPC workload optimization, RDMA/RoCE/InfiniBand protocols, Software-Defined Networking, QEMU/KVM virtualization, Deep learning frameworks (PyTorch, TensorFlow), Collective communication libraries (MPI, NCCL)
Technologies
Kubernetes, QEMU, KVM, InfiniBand, MPI, NCCL, PyTorch, TensorFlow
Responsibilities
Tune performance of GPU clusters and InfiniBand networks, Analyze and troubleshoot root causes of GPU/InfiniBand issues, Integrate new GPU hardware into infrastructure, Enhance automation systems for proactive monitoring, Configure and manage GPU devices and InfiniBand fabrics
Seniority
Senior, hands-on IC