HPC Software Development
Core
Installing, configuring, and maintaining High Performance Computing (HPC) systems, lab environments, and automation frameworks to support software development, testing, and validation.
Role type
Senior HPC Systems Engineer / DevOps
Builds
Reliable HPC clusters and automated management workflows for data-intensive workloads
Domain
High Performance Computing, Supercomputing, Linux System Administration
Deliverable
infrastructure
Required skills
Linux/Unix system administration, Python scripting, C/C++ programming, cluster management (Slurm/PBS/LSF), hardware troubleshooting, storage management (NFS/Lustre/GPFS), networking fundamentals, system monitoring (Prometheus/Grafana/Nagios), automation frameworks (Ansible)
Preferred skills
Parallel programming, Docker & Kubernetes, CI/CD, Infrastructure as Code, Infiniband networking, SQL
Technologies
Python, Ansible, Slurm, PBS, LSF, Red Hat, SUSE, Prometheus, Grafana, Nagios, NFS, Lustre, GPFS, Infiniband
Responsibilities
Install, configure, and maintain HPC systems including compute, storage, and management nodes; Manage HPC lab environments for software development and testing; Develop and maintain automation frameworks for system installation and monitoring; Troubleshoot hardware and system-level issues; Monitor system performance and resolve bottlenecks; Implement security, backup, and recovery best practices; Maintain system documentation and collaborate on system upgrades.
Seniority
Senior, hands-on IC