Senior HPCLM Optimization Engineer
Core
Optimize HPC infrastructure, job schedulers, and GPU accelerators to increase throughput, reduce contention, and improve user experience for EDA and HPC workloads.
Role type
Senior IC HPC infrastructure optimization engineer
Builds
High-performance compute clusters, job scheduling policies, and GPU-enabled infrastructure for engineering design teams
Domain
Semiconductor / High-Performance Computing / EDA
Deliverable
production ML models | infrastructure
Required skills
HPC environments, batch schedulers (LSF/Slurm/PBS), Linux systems, scripting/automation, telemetry analysis, workload profiling, performance analysis, GPU porting, GPU-aware scheduler configuration, license management (FlexLM)
Preferred skills
HyperScheduler, AWS/hybrid cloud scaling, capacity forecasting, AI/ML for anomaly detection, Python, SQL, Grafana, LLMs, AI-assisted automation
Technologies
LSF, Slurm, PBS, FlexLM, Python, SQL, Grafana, AWS
Responsibilities
Analyze cluster, queue, workload, and license telemetry to identify optimization opportunities; Drive scheduler and policy tuning for job placement and fairshare behavior; Profile critical EDA and HPC workloads to analyze CPU, GPU, memory, and I/O behavior; Partner with cross-functional teams to resolve systemic issues; Develop dashboards and KPIs for queue health and capacity forecasting; Apply automation and AI/ML-driven approaches for anomaly detection and congestion avoidance; Document best practices and operational playbooks
Seniority
Senior, hands-on IC