HPC Production Engineer
Core
Design, implement, maintain, and support high performance compute and storage systems for quantitative research data pipelines.
Role type
Senior HPC Production Engineer
Builds
High performance compute clusters, storage systems, and monitoring tooling for research teams
Domain
Financial technology / High Performance Computing
Deliverable
infrastructure
Required skills
Linux systems administration, HPC cluster management, parallel filesystems, batch scheduling systems, high-performance network interconnects, system configuration management, software development (Go/Python/C), performance profiling and debugging, root cause analysis
Preferred skills
Experience with Lustre, GPFS, Slurm, Grid Engine
Technologies
Linux, Go, Python, C, SaltStack, Ansible, Puppet, Lustre, GPFS, Slurm, Grid Engine
Responsibilities
Design and maintain HPC and storage systems; implement performance and fault monitoring; build tooling for software deployment at scale; collaborate with researchers to optimize infrastructure usage; manage vendor relationships and travel for meetings; participate in coordinated maintenance operations
Seniority
Senior, hands-on IC