HPC Engineer
Core
Design, operate, and optimize large-scale GPU compute environments for distributed model training and high-performance workloads.
Role type
HPC / GPU Cluster Engineer
Builds
Large-scale GPU clusters for distributed model training
Domain
Quantitative trading / High-performance computing
Deliverable
infrastructure
Required skills
RDMA programming, InfiniBand networking, Slurm scheduler management, high-performance shared storage administration, Linux system administration, Python scripting, Bash scripting, C/C++ programming
Preferred skills
GPUDirect RDMA, GPUDirect Storage, NCCL, CUDA drivers, OFED, RoCEv2, network topology design
Technologies
InfiniBand, RoCEv2, GPUDirect RDMA, GPUDirect Storage, NCCL, CUDA, OFED, Slurm, Lustre, BeeGFS, WEKA, VAST, DDN/ExaScaler
Responsibilities
Design and operate large-scale GPU clusters, optimize cluster performance and stability, manage storage layers and scheduling systems, develop tooling for monitoring and automation, troubleshoot network and system issues
Seniority
Mid-level, hands-on IC