Staff Site Reliability Engineer - Splunk
Core
Building a world-class, scalable Observability Platform using Splunk to enable SRE teams and business partners to monitor complex distributed systems.
Role type
Staff Site Reliability Engineer (Observability/Splunk)
Builds
Scalable observability infrastructure, automated agent deployments, and high-reliability log processing pipelines.
Domain
Cloud Infrastructure / Observability / Log Management
Deliverable
production ML models | infrastructure
Required skills
Splunk Cloud scaling (1000+ SVCs), Terraform, Go, Python, Ruby, Linux internals, Kubernetes/EKS, TCP/IP, DNS, Load Balancing, SPL, incident response, post-incident reviews, data-driven debugging
Preferred skills
OpenTelemetry, Vector, Splunk charge-back app, AWS/GCP native observability tools
Technologies
Splunk, Terraform, Go, Python, Ruby, Kubernetes, EKS, AWS, GCP, OpenTelemetry, Vector
Responsibilities
Design and maintain scalable observability infrastructure; optimize Splunk data collection, processing, and storage; lead incident response and post-incident reviews; automate deployment and scaling of observability agents.
Seniority
Staff, hands-on IC with strategic ownership