Senior Staff Site Reliability Engineer
Core
Senior technical individual contributor responsible for owning the technical strategy, responding to critical incidents, and preventing system failures for Splunk Cloud at global scale.
Role type
Senior Staff Site Reliability Engineer (IC)
Builds
Resilient automation frameworks, infrastructure architecture, and operational processes for Splunk Cloud SaaS platform.
Domain
Cloud Operations / Site Reliability Engineering / Enterprise SaaS
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Linux administration, cloud platform operations (AWS/GCP/Azure), scripting/automation (Python/Go), incident management, root cause analysis, technical leadership, cross-functional communication
Preferred skills
Splunk-specific knowledge (SPL, indexer clustering), observability platform expertise, enterprise-scale customer escalations, mentoring senior talent
Technologies
AWS, GCP, Azure, Python, Go/golang, Splunk, Linux
Responsibilities
Serve as escalation authority during P1/P2 critical incidents; own technical strategy for incident response and prevention; shape automation direction and infrastructure architecture; lead post-mortems and drive systemic improvements; mentor senior engineers and raise technical standards.
Seniority
Senior Staff, hands-on IC with strategic influence