CareerPlanGet AI match score →

High-Performance Computing DevOps Architect

2 Locations💼 Full-time🗓 2026-05-13 → 2026-07-30

Core

Architecting and optimizing high-performance computing (HPC) infrastructure for low-latency, high-throughput image processing and deep-learning workloads to support chip manufacturing process control equipment.

Role type

Senior IC HPC DevOps Architect

Builds

High-performance compute clusters, storage solutions, and networking infrastructure for semiconductor equipment

Domain

Semiconductor manufacturing / High-Performance Computing

Deliverable

infrastructure

Required skills

Linux system administration, HPC cluster computing, high-speed interconnects (InfiniBand, RoCE), parallel filesystems (Lustre, GPFS, BeeGFS), automation tools (Ansible, Chef, Salt-Stack), Python, Bash, storage solutions (RAID, IOPS tuning), networking concepts (IP addressing, routing, RDMA, VLAN), virtualization, containerization (Docker, Singularity), Software Defined Networking, remote management protocols (IPMI, Redfish), monitoring systems (Prometheus, Grafana), system profiling and tuning

Preferred skills

HPC workload orchestration (SLURM, Torque, LSF), deep-learning training/inference setup, private cloud infrastructure (Kubernetes, OpenStack, CloudStack), distributed HPC and parallel programming frameworks, low-latency data transfer technologies

Technologies

InfiniBand, RoCE, Lustre, GPFS, BeeGFS, Ansible, Chef, Salt-Stack, Python, Bash, Docker, Singularity, Prometheus, Grafana, SLURM, Torque, LSF, Kubernetes, OpenStack, CloudStack

Responsibilities

Design HPC infrastructure solutions including compute, networking, storage, and workload management; create and maintain system architecture diagrams and specifications; evaluate and select hardware and software components for HPC environments; install, configure, and maintain HPC systems; develop and implement automation scripts for system management and deployment; develop system benchmarks and profile systems to identify bottlenecks and optimize workflows; mitigate technical risks throughout the HPC development life cycle; ensure compute cluster resilience, reliability, and maintainability; stay abreast of latest HPC technologies

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗