Site Reliability Engineer, Compute Platform
Core
Ensure reliability of TikTok's major data warehouse products, services, and query engines (ClickHouse, Spark, Presto, Doris) while optimizing performance and managing incidents.
Role type
Senior IC Site Reliability Engineer (Compute Platform)
Builds
Data warehouse services, query engines, and underlying infrastructure for Big Data products
Domain
Big Data / Data Warehousing / Cloud Infrastructure
Required skills
Linux administration, computer networking, database management, Kubernetes, container orchestration, Python/Shell/Java/Go scripting, incident management, capacity planning, infrastructure automation
Preferred skills
Experience with ClickHouse, Hadoop, Doris, Spark, Presto, proactive performance optimization, cross-functional collaboration
Technologies
ClickHouse, Spark, Presto, Doris, Kubernetes, Hadoop, Linux, Python, Shell, Java, Go
Responsibilities
Maintain SLAs for data platform services, troubleshoot and resolve system outages, automate infrastructure provisioning and scaling, analyze performance patterns to prevent disruptions, lead incident postmortems, forecast infrastructure capacity needs, collaborate with development teams on reliability integration
Seniority
Senior, hands-on IC