Senior Site Reliability Engineer- (APJ-Remote)
Core
Building and leading processes to ensure the reliability, availability, scalability, and performance of ClickHouse Cloud infrastructure.
Role type
Senior Site Reliability Engineer (Cloud Infrastructure)
Builds
Scalable, secure, highly available, and fault-tolerant distributed systems for ClickHouse Cloud
Domain
Cloud computing, Distributed databases, Site Reliability Engineering
Required skills
Site Reliability Engineering, Go, Python, Cloud platforms (AWS/Azure/GCP), Distributed databases, SQL, Kubernetes, Docker Swarm, Ansible, Terraform, Puppet, Incident management, Chaos engineering, On-call management
Preferred skills
ClickHouse production experience, Production debugging, Data governance
Responsibilities
Design and implement scalable systems, Establish and manage SLOs/SLAs, Ensure monitoring and alerting for infrastructure components, Enhance incident response and post-mortem analysis, Drive Chaos initiatives, Manage on-call processes
Seniority
Senior, hands-on IC