IT Systems Engineer V
Core
Management and administration of the university's central research computing infrastructure, including supercomputing clusters, scalable storage systems, and specialized research software.
Role type
Senior IC HPC systems engineer
Builds
Production ML models | infrastructure
Domain
Higher education + High Performance Computing (HPC)
Deliverable
infrastructure
Required skills
Linux-based supercomputing cluster management, scalable parallel file systems (Lustre), research software administration, hardware specification and installation, job scheduler configuration (SLURM), performance tuning, secure enclave management
Preferred skills
CentOS/RedHat Enterprise Linux, Linux VM and container management, BASH/Python scripting, scientific computing tools (MPI, OpenMP, CUDA, VASP, Gaussian, Nextflow, Open OnDemand), Docker and Singularity administration, cluster management software (Warewulf, DUO), SELinux, Ceph, databases (Postgres, MySQL, MongoDB), NIST 800-171 compliance
Technologies
SLURM, Lustre, CentOS, RedHat Enterprise Linux, Ubuntu, Docker, Singularity, Warewulf, DUO, OpenHPC, Google Cloud, Open OnDemand
Responsibilities
Design, operate and manage highly complex university research supercomputing and storage systems; Manage a suite of specialized research software including licenses and OS support; Assist with specification, ordering and installation of new hardware, storage and networking; Perform problem analysis, monitor and tune performance, and configure job schedulers; Advise faculty and students on effective use of academic and supercomputing technology; Work with faculty to develop research grant proposals for computing resources; Consult on protected data handling and secure enclave management
Seniority
Senior, hands-on IC