CareerPlanSign in

Senior Cluster Site Reliability Engineer

Berkeley, CA💼 Full-time🗓 2026-04-17 → 2026-09-26

Core

Scale research compute clusters to support AI/ML research in finance, ensuring high uptime and reliability for on-prem and cloud infrastructure.

Role type

Senior Cluster Site Reliability Engineer (HPC/ML Infrastructure)

Builds

High-performance computing (HPC) platforms and batch compute systems for machine learning research

Domain

Finance, Artificial Intelligence, Machine Learning, High-Performance Computing

Deliverable

production ML models | infrastructure

Required skills

SRE/DevOps, HPC/batch compute frameworks, Kubernetes, Infrastructure-as-Code, Cloud infrastructure, Observability, Distributed storage, Scripting

Preferred skills

HPC frameworks (Slurm, Grid Engine), Kubernetes job orchestrators, Distributed computing frameworks, ML frameworks, Containerization, HPC networking, Security/IAM

Technologies

Slurm, Kueue, AWS Batch, GCP Batch, Kubeflow, MLflow, Horovod, Terraform, Ansible, AWS, GCP, Prometheus, Grafana, Loki, ELK, OpenTelemetry, Lustre, Ceph, S3, Docker, Podman, Singularity, InfiniBand, RDMA

Responsibilities

Triage and resolve urgent cluster outages; Define and track SLAs for cluster uptime; Diagnose systemic issues and engineer precision solutions; Develop custom observability mechanisms; Design policies for fair cluster usage; Forecast cluster growth and optimize cost/usability

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.