CareerPlanSign in

Senior Site Reliability Engineer (SRE) (Hybrid)

11 Locations💼 Full-time💰 $167,700–$167,700🗓 2026-08-07 → 2026-09-26

Core

Build, operate, and improve the reliability, scalability, and operational excellence of Splunk Agent Observability's deployment platform and production infrastructure for cloud and air-gapped customer deployments.

Role type

Senior Site Reliability Engineer (SRE)

Builds

Scalable, cost-effective evaluation and guardrails for AI agent resilience; deployment platforms and production infrastructure.

Domain

AI/ML Observability, Cloud Infrastructure, Kubernetes

Deliverable

production ML models | infrastructure

Required skills

Kubernetes production operations, CI/CD platform development, cloud platform experience (AWS/GCP), Python/Go programming, Infrastructure as Code (Terraform), database tuning, distributed systems debugging

Preferred skills

MLOps experience, observability/logging/alerting platforms, networking fundamentals (VPCs, DNS, routing, load balancing), air-gapped/on-prem environment deployment

Technologies

Kubernetes, Helm, Terraform, Python, Go, AWS, GCP

Responsibilities

Operate and improve Kubernetes-based production infrastructure; own customer deployments across cloud and air-gapped environments; build and improve deployment observability, monitoring, logging, and alerting; participate in production incident response and root cause analysis; design and develop internal tooling; debug complex production issues spanning Kubernetes, networking, storage, and application layers; collaborate with software engineers and customers to design secure, scalable, and reliable deployment architectures

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.