Senior Site Reliability Engineer - HPC
Core
Senior SRE owning end-to-end solutions for NVIDIA's global Compute Farm platform, ensuring uptime and QoS for HPC workloads across multi-cloud hybrid environments.
Role type
Senior Site Reliability Engineer (HPC)
Builds
Global services platform for High-Performance Computing (HPC) clusters
Domain
High-Performance Computing, Cloud Infrastructure, AI Infrastructure
Deliverable
production ML models | infrastructure
Required skills
HPC cluster management (Slurm, LSF, Kubernetes), Infrastructure as Code (IaC), CI/CD, Python/Go/Perl/Ruby scripting, observability, capacity planning, incident management, root cause analysis
Preferred skills
Open source maintenance, technical writing, AIOps/ML-driven operations
Technologies
Slurm, LSF, Kubernetes, AWS, GCP, OCI, Python, Go, Perl, Ruby
Responsibilities
Design and implement SRE solutions integrating with HPC schedulers and storage; automate provisioning via IaC; manage capacity and planning; detect performance issues; participate in on-call and incident reviews; mentor engineers and influence technical direction
Seniority
Senior, hands-on IC with mentorship