IT Systems Engineer V
Core
Management and administration of the university's central research computing infrastructure, including supercomputing clusters, software, data storage, and networking.
Role type
Senior IC HPC systems engineer
Builds
Production ML models | infrastructure
Domain
Higher education + High Performance Computing (HPC)
Deliverable
infrastructure
Required skills
Linux system administration, supercomputing clusters, parallel file systems (Lustre), job schedulers (SLURM), hardware specification and installation, research software administration, performance tuning, secure enclave management, Google Cloud Platform, scripting (BASH, Python), database administration (Postgres, MySQL, MongoDB)
Preferred skills
Physical layer networking, container management (Docker, Singularity), Ceph, scientific computing tools (MPI, OpenMP, CUDA, VASP, Gaussian, Nextflow, Open OnDemand), SELinux, CMMC 2.0, CUI, ITAR
Technologies
SLURM, Lustre, Google Cloud Platform, Docker, Singularity, CentOS, RedHat Enterprise Linux, Ubuntu, OpenHPC, Warewulf, DUO, Postgres, MySQL, MongoDB, MPI, OpenMP, CUDA, VASP, Gaussian, Nextflow, Open OnDemand, Ceph
Responsibilities
Provide technical support for design, operation and administration of highly complex university research supercomputing and storage systems; Manage a suite of specialized research software including licenses and OS support; Assist with specification, ordering and installation of new hardware, storage and associated networking; Perform problem analysis and resolution, monitor and tune performance, configure job schedulers to optimize operation; Advise and assist faculty, staff and graduate students on effective uses of academic and supercomputing technology; Work with faculty to develop research grant proposals to enhance research computing resources; Consult with users on handling protected data and support the management of the secure enclave on-prem and in Google cloud.
Seniority
Senior, hands-on IC