CareerPlanSign in

Senior HPC and LSF Operations Engineer

US, CA, Santa Clara💼 Full-time💰 $152,000–$152,000🗓 2026-09-24 → 2026-09-25

Core

Optimize, scale, and support workload scheduling systems (LSF, Slurm) for large-scale EDA and compute-intensive workloads to improve design velocity and infrastructure efficiency.

Role type

Senior HPC and LSF Operations Engineer

Builds

Reliable, observable, and automated EDA compute environments

Domain

Hardware Infrastructure / EDA Compute / High-Performance Computing

Deliverable

production ML models | infrastructure

Required skills

Linux systems administration (CentOS/RHEL), job scheduling systems (LSF, Slurm), problem solving under load, automation, SLO definition, documentation

Preferred skills

Reliability engineering practices, scheduler internals tuning, observability systems, container technologies (Docker, Singularity, Podman), cross-site standard adoption

Technologies

LSF, Slurm, CentOS, RHEL, Docker, Singularity, Podman

Responsibilities

Manage and optimize job scheduling systems in multi-site environments; Analyze performance data to identify bottlenecks and improve utilization; Lead problem solving across scheduler, OS, and workload layers; Implement automation to reduce manual effort; Define and track metrics and SLOs for service performance; Partner with customer teams to clarify requirements and drive issue closure

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.