Senior Systems HPC Engineer
Core
Building and optimizing large-scale GPU clusters for AI training and inference at the intersection of hardware and software.
Role type
Senior Systems HPC Engineer
Builds
Hyperscaler AI cloud platform components including GPU orchestration, storage, networking, and distributed communication layers.
Domain
Cloud Infrastructure / High-Performance Computing / AI Hardware
Deliverable
production ML models
Required skills
System-level software development, Linux administration and performance tuning, Server architecture knowledge, Performance-oriented programming (C/C++, Go, Python), GPU cluster troubleshooting, InfiniBand/RoCE networking, Virtualization (KVM/QEMU), Distributed communication (MPI, NCCL), Hardware qualification and acceptance testing.
Preferred skills
Experience with NVIDIA, Mellanox, or Intel hardware stacks.
Technologies
Linux, C/C++, Go, Python, InfiniBand, RoCE, KVM, QEMU, MPI, NCCL, PCIe, NVIDIA GPUs, Mellanox, Intel.
Responsibilities
Identify performance bottlenecks and drive improvements in cluster build, operation, tuning, and validation; Investigate and troubleshoot GPU cluster performance issues under real workloads; Evaluate and integrate new hardware, system configurations, and tuning approaches; Support complex performance-related escalations from internal teams and customers; Collaborate with infrastructure, software engineering, and hardware vendor teams; Contribute to hardware and cluster qualification to ensure performance expectations are met.
Seniority
Senior, hands-on IC