CareerPlanSign in

Lead Site Reliability Engineer - Ceph Storage

Montana - Remote🌐 Remote💼 Full-time🗓 2026-07-07 → 2026-09-26

Core

Design, architect, and operate one of the world's largest Ceph storage environments, serving GoDaddy's hosting infrastructure, internal services, OpenStack, and AI/HPC workloads.

Role type

Lead Senior Site Reliability Engineer (Storage Infrastructure)

Builds

Global Ceph clusters (object, block, file storage) powering hosting and AI/HPC platforms

Domain

Cloud Infrastructure / Distributed Storage Systems

Deliverable

production ML models | infrastructure

Required skills

Ceph architecture, distributed systems design, capacity planning, performance modeling, incident management, automation, observability, hardware selection, CRUSH topology, erasure coding

Preferred skills

OpenStack integration, large-scale storage modernization, cross-functional incident leadership

Technologies

Ceph, OpenStack, CRUSH, OSDs, S3/Swift, CephFS

Responsibilities

Design and architect large-scale production Ceph clusters; Lead fleet-wide capacity planning and hardware qualification; Own major platform upgrades and migrations; Resolve complex cross-functional production incidents; Establish automation and reliability practices

Seniority

Senior, hands-on IC with leadership responsibilities

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.