Database Reliability Engineer - Core Team
Core
Building and leading processes to ensure reliability, availability, scalability, and performance of ClickHouse Core, including incident response, post-mortem analysis, and chaos engineering.
Role type
Senior Site Reliability Engineer (Database Core)
Builds
ClickHouse Cloud infrastructure and core database services
Domain
Cloud computing + Distributed SQL databases
Required skills
Reliability Engineering, distributed database internals, SQL, production debugging, incident response, chaos engineering, scripting (Shell/Python), C++ code reading, cloud platforms (AWS/Azure/GCP)
Preferred skills
ClickHouse experience, bug fixing, root cause analysis
Technologies
ClickHouse, SQL, Shell, Python, C++, AWS, Azure, Google Cloud Platform
Responsibilities
Improve reliability and performance of ClickHouse core; Create metrics and alerts for production issues; Investigate customer problems and submit bug fixes; Manage incident response and post-mortem analysis; Drive chaos engineering initiatives; Manage on-call processes and escalation coordination
Seniority
Senior, hands-on IC
