CareerPlanGet AI match score →

Site Reliability Engineer - Ceph Storage

Austin🌐 Remote💼 Full-time🗓 2026-07-08 → 2026-07-31

Core

Senior Site Reliability Engineer owning the reliability, performance, scalability, and capacity of large-scale production Ceph environments supporting object, block, and file storage workloads.

Role type

Senior IC Site Reliability Engineer (Ceph Storage)

Builds

Object, block, and file storage platforms powering hosting, applications, internal infrastructure, and AI/HPC workloads

Domain

Cloud Infrastructure / Distributed Storage Systems

Deliverable

production ML models | infrastructure

Required skills

Linux internals, distributed storage systems, Ceph administration, Python, Shell scripting, configuration management (Ansible/SaltStack), incident response, root-cause analysis

Preferred skills

CephFS, RBD Mirroring, erasure coding, OpenStack integration, Kubernetes storage (CSI/Rook), Prometheus/Grafana, petabyte-scale migrations

Technologies

Ceph (RADOS, RGW, RBD, CephFS), Python, Shell, SaltStack, Ansible, Prometheus, Grafana, OpenStack, Kubernetes

Responsibilities

Diagnose and resolve complex distributed systems issues including recovery/backfill events and OSD instability; Design and build automation to reduce operational toil; Define and improve observability through SLIs, SLOs, and proactive alerting; Lead storage lifecycle initiatives including cluster expansions and hardware refreshes

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗