Staff Site Reliability Engineer - Observability GCP
Core
Building a world-class, scalable Observability Platform on Google Cloud to enable SRE teams and business partners to monitor complex distributed systems.
Role type
Staff Site Reliability Engineer (Observability)
Builds
Scalable observability infrastructure, automated agent deployment, and high-reliability data pipelines for Splunk and Grafana.
Domain
Cloud Infrastructure / Observability / SRE
Deliverable
infrastructure
Required skills
Google Cloud Platform (GCP) expertise, Terraform, Go/Python/Ruby, Kubernetes/GKE, Linux internals, TCP/IP networking, Splunk, Grafana, incident response, distributed systems debugging
Preferred skills
OpenTelemetry, Vector, Grafana Loki, AWS observability tools
Technologies
Terraform, Go, Python, Ruby, GKE, Splunk, Grafana, Linux, Kubernetes, OpenTelemetry, Vector, Loki, AWS
Responsibilities
Design and maintain scalable observability infrastructure using Terraform; optimize collection, processing, and storage of observability data; lead incident response and post-incident reviews; automate deployment and scaling of observability agents and collectors.
Seniority
Staff, hands-on IC with strategic ownership