Site Reliability Engineer - Ceph Storage
Core
Senior Site Reliability Engineer owning the reliability, performance, scalability, and capacity of large-scale production Ceph environments supporting object, block, and file storage workloads.
Role type
Senior IC Site Reliability Engineer (Ceph Storage)
Builds
Object, block, and file storage platforms powering hosting, applications, internal infrastructure, and AI/HPC workloads
Domain
Cloud Infrastructure / Distributed Storage Systems
Deliverable
production ML models | infrastructure
Required skills
Linux internals, distributed storage systems, Ceph administration, Python, Shell scripting, configuration management (Ansible/SaltStack), incident response, root-cause analysis
Preferred skills
CephFS, RBD Mirroring, erasure coding, OpenStack integration, Kubernetes storage (CSI/Rook), Prometheus/Grafana, petabyte-scale migrations
Technologies
Ceph (RADOS, RGW, RBD, CephFS), Python, Shell, SaltStack, Ansible, Prometheus, Grafana, OpenStack, Kubernetes
Responsibilities
Diagnose and resolve complex distributed systems issues including recovery/backfill events and OSD instability; Design and build automation to reduce operational toil; Define and improve observability through SLIs, SLOs, and proactive alerting; Lead storage lifecycle initiatives including cluster expansions and hardware refreshes
Seniority
Senior, hands-on IC