CareerPlanSign in

Senior HPC AI Cluster Engineer

US, CA, Santa Clara💼 Full-time💰 $176,000–$176,000🗓 2026-08-20 → 2026-09-25

Core

Design, implement, and maintain large-scale HPC/AI clusters, focusing on system design, tuning, and automation for at-scale compute runs.

Role type

Senior IC HPC/AI Cluster Infrastructure Engineer

Builds

Large-scale supercomputers and AI clusters

Domain

High-Performance Computing (HPC) and Artificial Intelligence (AI) Infrastructure

Deliverable

production ML models | infrastructure

Required skills

HPC and AI solution technologies, Linux internals and networking, job scheduling and orchestration, storage solutions, Python programming, automation and configuration management, virtual systems, cloud computing platforms

Preferred skills

CPU and/or GPU architecture, Kubernetes and container technologies, GPU-focused hardware/software, RDMA fabrics

Technologies

Slurm, K8s, Lustre, GPFS, Weka.io, InfiniBand, Ethernet, VMware, Hyper-V, KVM, Citrix, AWS, Azure, Google Cloud, Jenkins, Ansible, Puppet, chef, Cuda, DGX

Responsibilities

Design, implement, and maintain large scale HPC/AI clusters with monitoring, logging, and alerting; Manage Linux job/workload schedules and orchestration tools; Develop and maintain continuous integration and delivery pipelines; Develop tooling to automate deployment and management of large-scale infrastructure environments; Deploy monitoring solutions for servers, network, and storage; Perform troubleshooting from bare metal to application level; Support R&D activities and engage in POCs/POVs for future improvements

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.