Senior Site Reliability Engineering - Storage
Core
Own the reliability, performance, and scalability of global NAS, SAN, and Object Storage platforms powering critical internal and external services.
Role type
Senior Site Reliability Engineer (Storage)
Builds
Highly available, scalable storage infrastructure and automation for provisioning, monitoring, and lifecycle management.
Domain
Enterprise Storage Infrastructure (NAS, SAN, Object Storage)
Deliverable
production ML models | product features | infrastructure
Required skills
Enterprise storage platform design and operations, SRE practices (SLOs/SLIs, error budgets, incident management), Infrastructure as Code, container and virtualization platforms, scripting/programming, capacity planning and forecasting
Preferred skills
High-performance computing storage, AI/ML workload storage, large-scale data analytics storage, complex distributed system debugging, technical leadership and mentoring
Technologies
NAS, SAN, Object Storage, Terraform, Ansible, Puppet, SaltStack, Docker, Kubernetes, Python, Go, Shell
Responsibilities
Lead design, deployment, and operations of production storage platforms; architect storage solutions and drive end-to-end implementation; develop and improve automation for storage infrastructure; participate in on-call and incident response leading troubleshooting and root cause analysis; define and track SLOs/SLIs and error budgets; build and maintain runbooks and documentation; analyze capacity trends and recommend scaling strategies; mentor junior engineers and drive SRE adoption
Seniority
Senior, hands-on IC with mentorship responsibilities