Senior HPC AI Cluster Engineer
Core
Design, implement, and maintain large-scale HPC/AI clusters and infrastructure for supercomputing and AI workloads.
Role type
Senior HPC AI Cluster Engineer
Builds
Large-scale compute platforms, automated deployment tooling, and monitoring solutions for HPC/AI environments.
Domain
High-Performance Computing (HPC), Artificial Intelligence, GPU Computing, Data Center Infrastructure
Deliverable
infrastructure
Required skills
HPC and AI solution technologies, Linux networking and internals, job scheduling and orchestration, storage solutions, Python programming, automation and configuration management, virtual systems, cloud computing platforms
Preferred skills
CPU and/or GPU architecture, Kubernetes and container technologies, GPU-focused hardware/software, RDMA fabrics
Technologies
Slurm, K8s, Lustre, GPFS, Weka.io, InfiniBand, Ethernet, VMware, Hyper-V, KVM, Citrix, AWS, Azure, Google Cloud, Jenkins, Ansible, Puppet, Chef, Cuda, DGX
Responsibilities
Design and maintain large-scale HPC/AI clusters with monitoring and alerting; Manage Linux job/workload schedules and orchestration tools; Develop CI/CD pipelines; Develop tooling to automate deployment and management; Deploy monitoring solutions for servers, network, and storage; Perform troubleshooting from bare metal to application level; Develop and document standard methodologies; Support R&D activities and engage in POCs/POVs
Seniority
Senior, hands-on IC