Senior HPC and AI Cluster Administrator
Core
Design, deploy, and maintain large-scale HPC and AI clusters for US federal government clients, enabling GPU computing and scientific research workflows.
Role type
Senior HPC and AI Cluster Administrator
Builds
High-performance computing and AI infrastructure for defense, national security, and public safety organizations
Domain
US Federal Government / High-Performance Computing & AI Infrastructure
Deliverable
infrastructure
Required skills
HPC and AI solution technologies, job scheduling (Slurm, Kubernetes), Linux networking and internals, storage solutions (Lustre, GPFS, ZFS, XFS), automation (Python, Bash, Gitops), networking protocols (InfiniBand, Ethernet), private cloud platforms (VMware, Hyper-V, KVM, OpenShift, Nutanix), public cloud platforms (AWS, Azure)
Preferred skills
GPU architectures and MIG, container orchestration (Docker), AI workflow technologies (Apache Airflow, Prefect, Dagster), RDMA fabrics, compliance (DISA STIG, CIS), NVIDIA certifications, VMware certifications
Technologies
Kubernetes, Slurm, Python, Bash, Gitops, InfiniBand, Ethernet, Lustre, GPFS, ZFS, XFS, VMware, Hyper-V, KVM, OpenShift, Nutanix, AWS, Azure, Apache Airflow, Prefect, Dagster, Docker
Responsibilities
Design, deploy, and maintain HPC/AI clusters; Manage AI jobs workflows using scheduling technology; Support and maintain continuous integration and delivery pipelines; Troubleshoot and fix issues from bare metal to application level; Support Research, Development, and Operational activities
Seniority
Senior, hands-on IC