Senior HPC AI Cluster Engineer
Core
Design, implement, and maintain large-scale HPC/AI clusters, focusing on system design, tuning, and automation for at-scale compute runs.
Role type
Senior IC HPC/AI Cluster Infrastructure Engineer
Builds
Large-scale supercomputers and AI clusters
Domain
High-Performance Computing (HPC) and Artificial Intelligence (AI) Infrastructure
Deliverable
production ML models | infrastructure
Required skills
HPC and AI solution technologies, Linux internals and networking, job scheduling and orchestration, storage solutions, Python programming, automation and configuration management, virtual systems, cloud computing platforms
Preferred skills
CPU and/or GPU architecture, Kubernetes and container technologies, GPU-focused hardware/software, RDMA fabrics
Technologies
Slurm, K8s, Lustre, GPFS, Weka.io, InfiniBand, Ethernet, VMware, Hyper-V, KVM, Citrix, AWS, Azure, Google Cloud, Jenkins, Ansible, Puppet, chef, Cuda, DGX
Responsibilities
Design, implement, and maintain large scale HPC/AI clusters with monitoring, logging, and alerting; Manage Linux job/workload schedules and orchestration tools; Develop and maintain continuous integration and delivery pipelines; Develop tooling to automate deployment and management of large-scale infrastructure environments; Deploy monitoring solutions for servers, network, and storage; Perform troubleshooting from bare metal to application level; Support R&D activities and engage in POCs/POVs for future improvements
Seniority
Senior, hands-on IC