Site Reliability Engineer III
Core
Senior technical leader for HPC Enablement, defining operational standards, leading multi-team delivery, and translating researcher needs into scalable compute enablement designs.
Role type
Senior IC Site Reliability Engineer (HPC Enablement)
Builds
Scalable HPC infrastructure, compute enablement patterns, and operational guardrails for scientific workloads.
Domain
High-Performance Computing (HPC) / Scientific Computing
Deliverable
production ML models | infrastructure
Required skills
HPC enablement at scale, systems design at scale, performance tuning, SLO management, incident response, mentorship, stakeholder influence
Preferred skills
Kubernetes at scale, CI/CD, observability, security-by-design, standard-setting
Technologies
Kubernetes, CI/CD, HPC schedulers
Responsibilities
Own the compute reliability and enablement roadmap, define onboarding playbooks and golden paths, optimize scheduler configuration and resource allocation, conduct workload profiling and performance tuning, lead incident response, mentor engineers, partner with scientific teams to translate requirements into infrastructure patterns
Seniority
Senior, hands-on IC with strategy & mentorship