Senior Site Reliability Engineer
Core
Infrastructure and reliability engineer for the Data Replication team, ensuring high availability and performance for over 3 million weekly sync jobs across multiple regions and clouds.
Role type
Senior Site Reliability Engineer (Infrastructure & Platform)
Builds
Kubernetes clusters, CI/CD pipelines, secrets management, networking, and cloud resource configuration for the Data Replication platform.
Domain
Cloud Infrastructure, Data Engineering, Platform Engineering
Deliverable
infrastructure
Required skills
Kubernetes, Helm, Terraform, AWS, GCP, Prometheus, Grafana, Datadog, CI/CD pipeline ownership, backend code instrumentation, AI/LLM tooling for automation
Preferred skills
Data pipelines, replication systems, ETL/ELT platforms, Control plane / data plane architectures, Internal developer platforms, Airbyte, CDKs, connector-based architectures
Technologies
Kubernetes, Helm, Terraform, AWS, GCP, Prometheus, Grafana, Datadog
Responsibilities
Own infrastructure underpinning the Data Replication platform; Partner with product engineers to integrate features with infrastructure; Maintain and enhance observability, alerting, and anomaly detection; Maintain and enhance AI-augmented release and internal tooling; Set the infrastructure bar by building self-serve tooling and coaching engineers.
Seniority
Senior, hands-on IC