CareerPlanSign in

Senior Site Reliability Engineer - Storage

San Francisco Office (Fremont St)💼 Full-time🗓 2026-08-10 → 2026-09-25

Core

Own the reliability, performance, and capacity health of production storage fleets across data centers for an AI cloud infrastructure platform.

Role type

Senior Site Reliability Engineer (Storage)

Builds

Production storage systems, monitoring dashboards, and self-healing automation for AI compute workloads.

Domain

AI Cloud Infrastructure / Distributed Storage Systems

Deliverable

production ML models | infrastructure

Required skills

Linux systems operations, Software-Defined Storage (SDS), incident response, monitoring and logging, Kubernetes, CI/CD, Infrastructure as Code, storage protocols

Preferred skills

VAST or Weka, Enterprise storage (NetApp, Dell PowerScale), Kubernetes CSI drivers, SR-IOV/virtualization, GPUDirect Storage/RDMA/InfiniBand

Technologies

CEPH, Lustre, GPFS, Prometheus, Grafana, Alertmanager, Datadog, SumoLogic, Kubernetes, ArgoCD, Helm, Kustomize, Docker, Podman, Python, Go, Terraform, Ansible, Jenkins, GitHub Actions, BuildKite, NFS, SMB, S3, NVMe-oF, TCP, Vector DB, SQL

Responsibilities

Operate and maintain storage fleet health across data centers; build monitoring and alerting systems; investigate and resolve storage incidents; automate incident-response workflows; design self-healing automation for failure modes; implement CI/CD pipelines for storage tooling; partner on software-defined storage deployment; diagnose low-level I/O and network issues; participate in on-call rotation.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.