Lead Site Reliability Engineer - Ceph Storage
Core
Design, architect, and operate one of the world's largest Ceph storage environments, serving GoDaddy's hosting infrastructure, internal services, OpenStack, and AI/HPC workloads.
Role type
Lead Senior Site Reliability Engineer (Storage Infrastructure)
Builds
Global Ceph clusters (object, block, file storage) powering hosting and AI/HPC platforms
Domain
Cloud Infrastructure / Distributed Storage Systems
Deliverable
production ML models | infrastructure
Required skills
Ceph architecture, distributed systems design, capacity planning, performance modeling, incident management, automation, observability, hardware selection, CRUSH topology, erasure coding
Preferred skills
OpenStack integration, large-scale storage modernization, cross-functional incident leadership
Technologies
Ceph, OpenStack, CRUSH, OSDs, S3/Swift, CephFS
Responsibilities
Design and architect large-scale production Ceph clusters; Lead fleet-wide capacity planning and hardware qualification; Own major platform upgrades and migrations; Resolve complex cross-functional production incidents; Establish automation and reliability practices
Seniority
Senior, hands-on IC with leadership responsibilities