Senior HPC and LSF Operations Engineer
Core
Optimize, scale, and support workload scheduling systems (LSF, Slurm) for large-scale EDA and compute-intensive workloads to improve design velocity and infrastructure efficiency.
Role type
Senior HPC and LSF Operations Engineer
Builds
Reliable, observable, and automated EDA compute environments
Domain
Hardware Infrastructure / EDA Compute / High-Performance Computing
Deliverable
production ML models | infrastructure
Required skills
Linux systems administration (CentOS/RHEL), job scheduling systems (LSF, Slurm), problem solving under load, automation, SLO definition, documentation
Preferred skills
Reliability engineering practices, scheduler internals tuning, observability systems, container technologies (Docker, Singularity, Podman), cross-site standard adoption
Technologies
LSF, Slurm, CentOS, RHEL, Docker, Singularity, Podman
Responsibilities
Manage and optimize job scheduling systems in multi-site environments; Analyze performance data to identify bottlenecks and improve utilization; Lead problem solving across scheduler, OS, and workload layers; Implement automation to reduce manual effort; Define and track metrics and SLOs for service performance; Partner with customer teams to clarify requirements and drive issue closure
Seniority
Senior, hands-on IC