Site Reliability Engineer - ClickHouse
Core
Operating and scaling a petabyte-scale, self-managed ClickHouse data warehouse on AWS to support product analytics and AI features.
Role type
Senior Site Reliability Engineer (ClickHouse/OLAP)
Builds
A high-performance, automated data platform for product analytics and AI agents.
Domain
Cloud Infrastructure (AWS) + OLAP Databases (ClickHouse)
Deliverable
production ML models | infrastructure
Required skills
ClickHouse internals, AWS EC2/VM management, Terraform, Ansible, Linux systems administration, stateful system operations, performance debugging, incident response
Preferred skills
Experience with petabyte-scale data workloads, designing self-healing automation, deep ownership of production systems
Technologies
ClickHouse, AWS (EC2, S3, VPC), Terraform, Ansible, Linux
Responsibilities
Managing large fleets of EC2-based VMs and networking for data-intensive workloads; Improving operational tooling for deploys, schema changes, backups, and restores; Collaborating with database engineers to translate needs into infra solutions; Reducing operational load via automation and self-healing patterns; Participating in on-call and incident response
Seniority
Senior, hands-on IC