CareerPlanSign in

Senior Site Reliability Engineer - HPC

3 Locations💼 Full-time💰 $152,000–$152,000🗓 2026-09-21 → 2026-09-25

Core

Senior SRE owning end-to-end solutions for NVIDIA's global Compute Farm platform, ensuring uptime and QoS for HPC workloads across multi-cloud hybrid environments.

Role type

Senior Site Reliability Engineer (HPC)

Builds

Global services platform for High-Performance Computing (HPC) clusters

Domain

High-Performance Computing, Cloud Infrastructure, AI Infrastructure

Deliverable

production ML models | infrastructure

Required skills

HPC cluster management (Slurm, LSF, Kubernetes), Infrastructure as Code (IaC), CI/CD, Python/Go/Perl/Ruby scripting, observability, capacity planning, incident management, root cause analysis

Preferred skills

Open source maintenance, technical writing, AIOps/ML-driven operations

Technologies

Slurm, LSF, Kubernetes, AWS, GCP, OCI, Python, Go, Perl, Ruby

Responsibilities

Design and implement SRE solutions integrating with HPC schedulers and storage; automate provisioning via IaC; manage capacity and planning; detect performance issues; participate in on-call and incident reviews; mentor engineers and influence technical direction

Seniority

Senior, hands-on IC with mentorship

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.