HPC Engineer
Core
Operate, support, and modernize high-performance compute (HPC) platforms and job scheduling services to enable engineering teams to design and verify products.
Role type
Senior HPC Operations Engineer (SRE)
Builds
Production HPC environments, automation frameworks, and cloud-native scheduling services for engineering workloads.
Domain
Semiconductor engineering / High-Performance Computing
Deliverable
production ML models | infrastructure
Required skills
HPC environment operations, job scheduler administration (LSF, Slurm, PBS), Linux system administration, Python/Bash scripting, incident management and root cause analysis, observability (Prometheus, Grafana), CI/CD pipeline management, public/private/hybrid cloud platforms, Kubernetes, Infrastructure as Code.
Preferred skills
EDA tool familiarity, license-aware scheduling, container orchestration (Docker, Kubernetes), Terraform, Ansible, AI-assisted tooling for diagnostics.
Technologies
IBM Spectrum LSF, Slurm, PBS, Grid Engine, RHEL, Python, Bash, Dynatrace, Prometheus, Grafana, AWS, GCP, Azure, OpenStack, Kubernetes, Docker, Terraform, Ansible, Jira, Confluence.
Responsibilities
Operate and improve HPC platforms and job schedulers; develop automation and self-service capabilities; support production environments including incident response and recovery; collaborate with engineering users to optimize workload performance and resource utilization; contribute to cloud HPC integration and modernization initiatives; define standards for DevOps, SRE, and infrastructure automation.
Seniority
Senior, hands-on IC