HPC Operations Engineer
Core
Provide front-line operational support for 24/7 Linux HPC compute, storage, and interconnects to support data pipelines and quantitative research.
Role type
HPC Operations Engineer
Builds
Production computing environments for quantitative research
Domain
High Performance Computing / Financial Technology
Deliverable
infrastructure
Required skills
Linux systems administration, root cause analysis, scripting (Go, Python, C), vendor management, performance monitoring, fault monitoring, documentation development
Preferred skills
HPC environments (parallel filesystems like Lustre/GPFS, batch systems like Slurm/Grid Engine), high-performance network interconnects (RDMA), multi-vendor hardware experience
Technologies
Linux, RDMA fabrics, parallel filesystems, HPC batch schedulers, FUSE filesystems, internal Jump software, multi-vendor hardware
Responsibilities
Solve problem reports and manage the entire problem lifecycle, respond to alerts in a timely fashion, participate in large coordinated maintenance operations, write code for diagnosing and automating tasks, collaborate on testing infrastructures, manage relationships with outside vendors, implement and support performance and fault monitoring systems, develop and monitor tools for maintaining the production environment
Seniority
Mid-level, hands-on IC