HPC Operations Engineer
Core
Architecting and optimizing a private compute cloud for chip development, ensuring system reliability and efficiency from bare metal to application level.
Role type
Senior IC HPC Operations Engineer
Builds
Private compute cloud infrastructure for next-generation chip development
Domain
Semiconductor / High-Performance Computing
Deliverable
infrastructure
Required skills
Linux distribution management (CentOS/RHEL), Python scripting, Bash scripting, Ansible, Docker, cluster configuration management, problem-solving
Preferred skills
NFS, automounter, LDAP, DNS, TCP/IP networking, job scheduler administration (LSF/SLURM), FlexLM license management, Perl, InfiniBand, RDMA, RoCE, Lustre, GPFS
Responsibilities
Troubleshoot support requests in large-scale HPC environments, improve deployment automation and observability, ensure accurate OS and configuration on compute servers, resolve complex issues from bare metal to application level, collaborate with specialist teams to drive issue closure, work with domain experts to optimize chip development processes
Seniority
Mid-level, hands-on IC