Senior Site Reliability Engineer - Storage
Core
Own the reliability, performance, and capacity health of production storage fleets across data centers for an AI cloud infrastructure platform.
Role type
Senior Site Reliability Engineer (Storage)
Builds
Production storage systems, monitoring dashboards, and self-healing automation for AI compute workloads.
Domain
AI Cloud Infrastructure / Distributed Storage Systems
Deliverable
production ML models | infrastructure
Required skills
Linux systems operations, Software-Defined Storage (SDS), incident response, monitoring and logging, Kubernetes, CI/CD, Infrastructure as Code, storage protocols
Preferred skills
VAST or Weka, Enterprise storage (NetApp, Dell PowerScale), Kubernetes CSI drivers, SR-IOV/virtualization, GPUDirect Storage/RDMA/InfiniBand
Technologies
CEPH, Lustre, GPFS, Prometheus, Grafana, Alertmanager, Datadog, SumoLogic, Kubernetes, ArgoCD, Helm, Kustomize, Docker, Podman, Python, Go, Terraform, Ansible, Jenkins, GitHub Actions, BuildKite, NFS, SMB, S3, NVMe-oF, TCP, Vector DB, SQL
Responsibilities
Operate and maintain storage fleet health across data centers; build monitoring and alerting systems; investigate and resolve storage incidents; automate incident-response workflows; design self-healing automation for failure modes; implement CI/CD pipelines for storage tooling; partner on software-defined storage deployment; diagnose low-level I/O and network issues; participate in on-call rotation.
Seniority
Senior, hands-on IC