CareerPlanGet AI match score →

Senior Cluster Site Reliability Engineer

Berkeley, CA, US💼 Full-time💰 $205,000–$235,000🗓 2026-06-04 → 2026-07-29

Core

Scale research compute clusters to support AI/ML research in finance, ensuring high uptime and reliability for technical staff.

Role type

Senior Site Reliability Engineer (HPC/Cluster Operations)

Builds

World-class HPC platform for researchers to solve cutting-edge machine learning problems at scale

Domain

Finance / High-Performance Computing / Machine Learning Infrastructure

Deliverable

infrastructure

Required skills

SRE/DevOps, HPC/batch compute frameworks (Slurm, Kueue), cloud infrastructure (AWS/GCP), infrastructure-as-code (Terraform, Ansible), observability stacks (Prometheus, Grafana, Loki, ELK, OpenTelemetry), distributed storage (Lustre, Ceph, S3), scripting (Python, Ruby)

Preferred skills

Kubernetes-based job orchestrators (Airflow, Kueue, Kubeflow Pipelines), distributed computing frameworks (Ray, Modin, Dask, Spark), ML frameworks (PyTorch, Tensorflow, JAX, Horovod, DeepSpeed), containerization (Docker, Podman, Singularity), HPC networking (InfiniBand, RDMA), security/IAM foundations

Technologies

Slurm, Kueue, AWS, GCP, Terraform, Ansible, Prometheus, Grafana, Loki, ELK, OpenTelemetry, Lustre, Ceph, S3, Python, Ruby, Kubernetes, Airflow, Kubeflow Pipelines, Ray, Modin, Dask, Spark, PyTorch, Tensorflow, JAX, Horovod, DeepSpeed, Docker, Podman, Singularity, InfiniBand, RDMA

Responsibilities

Be a first responder for cluster outages and triage urgent issues; ensure high cluster uptime and define/track SLAs; diagnose systemic problems and engineer precision solutions; develop robust metrics and custom observability mechanisms; help design policies and enforcement mechanisms for fair cluster usage; assist in forecasting cluster growth and optimizing operations for cost and usability

Seniority

Senior, hands-on IC

Sourced via efinancialcareers · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on eFinancialCareers ↗