CareerPlanSign in

Site Reliability Engineer II

💼 Full-time🗓 2026-08-11 → 2026-09-25

Core

Build automation, maintain observability, and support incident response to ensure the stability, scalability, and reliability of customer-facing cloud storage services.

Role type

Site Reliability Engineer II (hands-on IC)

Builds

Cloud storage infrastructure and operational tooling for 500K+ customers

Domain

Cloud storage / Distributed systems

Deliverable

production ML models | product features | dashboards & analysis | infrastructure

Required skills

Linux systems administration, scripting (Python/Bash/Go), container orchestration (Kubernetes/Docker), incident response, root cause analysis, CI/CD pipelines, infrastructure as code (Terraform/Ansible)

Preferred skills

SaaS/service provider experience, ITIL/OSS practices, SLO/SLA management, cloud platforms (AWS/GCP/Azure)

Technologies

Prometheus, Grafana, Catchpoint, ELK, Terraform, Ansible, Jenkins, Kubernetes, Docker, AWS, GCP, Azure

Responsibilities

Monitor service health using SLIs/SLOs and error budgets, participate in on-call rotations and incident response, develop automation to reduce manual toil, contribute to monitoring and logging frameworks, assist in capacity planning and disaster recovery exercises, document systems and operational playbooks

Seniority

Mid-level, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.