Senior HPC DevOps Engineer, NCS
Core
Design, implement, and maintain large-scale HPC/AI clusters with state-of-the-art monitoring, logging, and alerting systems.
Role type
Senior HPC DevOps Engineer
Builds
Large-scale HPC/AI clusters, performance platforms, and automated deployment pipelines
Domain
High-Performance Computing, Artificial Intelligence, GPU Computing
Deliverable
production ML models | infrastructure
Required skills
Infrastructure as Code (IaC), CI/CD pipeline development, automation scripting, complex networking, troubleshooting from bare metal to application level, technical leadership, R&D support
Preferred skills
CPU/GPU architecture knowledge, job scheduling and orchestration (Slurm, Kubernetes), professional networking training
Technologies
Jenkins, Ansible, Puppet/Chef, Kubernetes, Apache Kafka, Lustre, GPFS, ZFS, XFS, VMware, Hyper-V, KVM, Citrix, AWS, Azure, Google Cloud