Ingénieur.e administration système HPC - Industrie - Le Plessis Robinson
Core
Maintaining and operating the world's largest private HPC cluster for scientific simulation in the energy sector.
Role type
Senior HPC Systems Administrator
Builds
High-performance computing infrastructure and simulation environments for energy industry clients
Domain
High-Performance Computing (HPC) / Scientific Simulation / Energy
Deliverable
infrastructure
Required skills
Linux administration, HPC cluster management, network administration (Ethernet, InfiniBand, Lustre), configuration management (Puppet, Ansible), monitoring (Centreon, Grafana, Prometheus), scheduling (Slurm, PBS, LSF), parallel storage (Lustre, GPFS), Python/Perl/Bash scripting
Preferred skills
Hardware troubleshooting, English fluency
Technologies
RedHat, Python, Perl, Bash, Ethernet, InfiniBand, Lustre, Puppet, Ansible, Centreon, Grafana, Prometheus, Slurm, PBS, LSF, GPFS
Responsibilities
Install, upgrade, and configure HPC middleware software; evolve system and infrastructure architectures to integrate new hardware; ensure maintenance, correction, and support (MCO/MCS) of calculation clusters; monitor and follow up on simulation production running on clusters
Seniority
Senior, hands-on IC