Senior Site Reliability Engineer- Remote
Core
Building and leading processes to ensure the reliability, availability, scalability, and performance of ClickHouse Cloud infrastructure.
Role type
Senior Site Reliability Engineer
Builds
Elastic, limitless scale, high-performance ClickHouse Cloud
Domain
Cloud infrastructure, distributed databases, real-time analytics
Deliverable
production ML models | infrastructure
Required skills
Go, Python, cloud computing platforms (AWS, Azure, GCP), distributed databases, SQL, container orchestration (Kubernetes, Docker Swarm), automation and configuration management (Ansible, Terraform, Puppet), production debugging, incident management, SLO/SLA management, chaos engineering
Preferred skills
ClickHouse expertise, efficiency and data governance focus
Technologies
ClickHouse, Kubernetes, Docker Swarm, Ansible, Terraform, Puppet, AWS, Azure, Google Cloud Platform
Responsibilities
Design and implement scalable, secure, and highly available systems; Establish and manage SLOs and SLAs; Ensure monitoring and alerting for infrastructure components; Enhance incident response processes and post-mortem analysis; Continuously improve reliability and performance; Plan and drive Chaos initiatives; Manage on-call processes and escalation best practices
Seniority
Senior, hands-on IC