Senior Site Reliability Engineer, Data Infrastructure
Core
Own the reliability and performance of a Kubernetes-based data platform powering data ingestion, transformation, analytics, and internal AI workloads.
Role type
Senior Site Reliability Engineer (Data Infrastructure)
Builds
Highly available, multi-region systems for data platform operations
Domain
Cloud-native infrastructure, Kubernetes, Data Infrastructure
Deliverable
infrastructure
Required skills
Kubernetes cluster design and operations, CI/CD system ownership, incident response and SLO management, geo-replicated active-active system design, observability stack implementation, infrastructure as code, distributed system performance tuning, cloud-native security practices
Preferred skills
Data platform operations (Spark, Airflow, Kafka, Flink), service mesh technologies, regulated environment compliance, internal developer platform building
Technologies
Kubernetes, Argo CD, GitHub Actions, Prometheus, Grafana, OpenTelemetry, Helm, Terraform, Pulumi, Istio, Linkerd
Responsibilities
Design and operate highly available, multi-region systems; scale infrastructure and improve deployment pipelines; harden security posture; evolve DevSecOps practices; define SLI/SLO/SLA and manage error budgets; respond to incidents and conduct postmortems; tune system performance and plan capacity
Seniority
Senior, hands-on IC