Site Reliability Engineer II
Core
Build automation, maintain observability, and support incident response to ensure the stability, scalability, and reliability of customer-facing cloud storage services.
Role type
Site Reliability Engineer II (hands-on IC)
Builds
Cloud storage infrastructure and operational tooling for 500K+ customers
Domain
Cloud storage / Distributed systems
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Linux systems administration, scripting (Python/Bash/Go), container orchestration (Kubernetes/Docker), incident response, root cause analysis, CI/CD pipelines, infrastructure as code (Terraform/Ansible)
Preferred skills
SaaS/service provider experience, ITIL/OSS practices, SLO/SLA management, cloud platforms (AWS/GCP/Azure)
Technologies
Prometheus, Grafana, Catchpoint, ELK, Terraform, Ansible, Jenkins, Kubernetes, Docker, AWS, GCP, Azure
Responsibilities
Monitor service health using SLIs/SLOs and error budgets, participate in on-call rotations and incident response, develop automation to reduce manual toil, contribute to monitoring and logging frameworks, assist in capacity planning and disaster recovery exercises, document systems and operational playbooks
Seniority
Mid-level, hands-on IC