HPC Systems Engineer
Core
Installs, configures, and maintains large computer clusters/servers and software for high-end research computing resources.
Role type
HPC Systems Engineer
Builds
Centrally managed High-Performance Computing (HPC), storage, and visualization resources for researchers.
Domain
Academic research computing / High-Performance Computing
Deliverable
infrastructure
Required skills
Linux build automation, shell scripting, job management tools, network storage subsystems, distributed file systems, scientific app software configuration, performance monitoring
Technologies
Puppet, Ansible, Git, Docker, Python, Shell, Perl, SLURM, Moab, TORQUE, PBS, XCAT, ROCKS, IBM NetApp, Data Direct Network, LSI, GPFS, Lustre, Gluster
Responsibilities
Install, configure, and maintain large computer clusters/servers; manage system network switches, parallel file systems, and HPC software stacks; diagnose and resolve system operational problems; coordinate with vendors for hardware/software issues; assist users with access and help desk tickets; build and deploy open source and vendor software; provide reliable backups/restores; maintain system security; document system administration procedures.
Seniority
Mid-level, hands-on IC