CareerPlanGet AI match score →

Senior Systems HPC Engineer

United States💼 Full-time🗓 2026-07-14 → 2026-07-31

Core

Building and optimizing large-scale GPU clusters for AI training and inference at the intersection of hardware and software.

Role type

Senior Systems HPC Engineer

Builds

Hyperscaler AI cloud platform components including GPU orchestration, storage, networking, and distributed communication layers.

Domain

Cloud Infrastructure / High-Performance Computing / AI Hardware

Deliverable

production ML models

Required skills

System-level software development, Linux administration and performance tuning, Server architecture knowledge, Performance-oriented programming (C/C++, Go, Python), GPU cluster troubleshooting, InfiniBand/RoCE networking, Virtualization (KVM/QEMU), Distributed communication (MPI, NCCL), Hardware qualification and acceptance testing.

Preferred skills

Experience with NVIDIA, Mellanox, or Intel hardware stacks.

Technologies

Linux, C/C++, Go, Python, InfiniBand, RoCE, KVM, QEMU, MPI, NCCL, PCIe, NVIDIA GPUs, Mellanox, Intel.

Responsibilities

Identify performance bottlenecks and drive improvements in cluster build, operation, tuning, and validation; Investigate and troubleshoot GPU cluster performance issues under real workloads; Evaluate and integrate new hardware, system configurations, and tuning approaches; Support complex performance-related escalations from internal teams and customers; Collaborate with infrastructure, software engineering, and hardware vendor teams; Contribute to hardware and cluster qualification to ensure performance expectations are met.

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗