HPC Operations Engineer
Core
Provide front-line operational support for 24/7 Linux HPC compute, storage, and interconnects, solving problem reports and responding to alerts for Jump's research community.
Role type
Senior IC HPC Operations Engineer
Builds
Production Linux HPC environments including RDMA fabrics, parallel filesystems, and batch schedulers for global financial research
Domain
High Performance Computing (HPC) / Financial Technology
Deliverable
infrastructure
Required skills
Linux systems administration, root cause analysis, programming/scripting (Go, Python, C), vendor management, performance monitoring, fault monitoring, documentation development
Preferred skills
High performance computing (HPC), parallel filesystems (Lustre, GPFS), batch systems (Slurm, Grid Engine), high-performance network interconnects
Technologies
Linux, RDMA, Lustre, GPFS, Slurm, Grid Engine, FUSE, Go, Python, C
Responsibilities
Provide front-line operational support for 24/7 Linux HPC compute, storage, and interconnects; Solve problem reports and questions posed by members of Jump's research community; Respond to alerts in a timely fashion; Participate in large, coordinated maintenance operations; Write code for diagnosing, resolving, and triaging difficult problems; Manage relationships with outside vendors; Implement and support performance monitoring and fault monitoring systems; Develop and improve systems and user documentation
Seniority
Senior, hands-on IC