HPC Systems Administrator
Core
Designing, deploying, and maintaining secure, large-scale High-Performance Computing (HPC) infrastructure for university research, including CPU/GPU clusters, storage, and network systems.
Role type
Senior HPC Systems Security Engineer
Builds
Scalable HPC clusters, storage infrastructure, and secure operational environments for research workloads
Domain
Academic research computing / High-Performance Computing
Deliverable
production ML models | infrastructure
Required skills
Linux system administration (RHEL/Rocky), GPU infrastructure tuning and management, job scheduler administration (Slurm, Torque, PBS, LSF), infrastructure automation (Ansible, Puppet, Chef, Salt), distributed storage systems (Lustre, Ceph, Gluster), Python or Bash scripting, vulnerability scanning and patch management, InfiniBand concepts
Preferred skills
Security compliance (NIST 800-53, 800-171, 800-223, FIPS), monitoring tool implementation (CheckMK, Zabbix, Prometheus, Grafana), documentation of security procedures, translating scientific goals into computational requirements
Technologies
Slurm, Torque, PBS, LSF, Ansible, Puppet, Chef, Salt, xCAT, Confluent, Warewulf, CheckMK, Zabbix, Nagios, Prometheus, Grafana, Storage Scale, Lustre, Gluster, BeeGFS, Ceph, InfiniBand, RHEL, Rocky
Responsibilities
Design and deploy CPU/GPU HPC clusters and storage infrastructure; tune GPU nodes for optimal performance; develop monitoring and observability tools; enforce security procedures and compliance; manage job scheduling environments; troubleshoot hardware and software stack issues; implement backup and disaster recovery capabilities; perform vulnerability scanning and patch management; maintain system documentation and inventory tracking
Seniority
Senior, hands-on IC