CareerPlanSign in

Staff Site Reliability Engineer (SRE) (Hybrid)

11 Locations💼 Full-time💰 $186,900–$186,900🗓 2026-08-07 → 2026-09-26

Core

Provide technical leadership for the reliability, scalability, and operational architecture of Splunk Agent Observability's platform to ensure AI agents behave as intended.

Role type

Staff Site Reliability Engineer (SRE)

Builds

Scalable, cost-effective evaluation and guardrails for AI agents; deployment platforms for cloud and air-gapped environments.

Domain

AI resilience, observability, cloud infrastructure, and distributed systems.

Deliverable

production ML models | infrastructure

Required skills

Site Reliability Engineering, Kubernetes, distributed systems design, cloud platforms (AWS/GCP), CI/CD automation, incident response, technical leadership, Python, Go, Infrastructure as Code (Terraform), networking, databases, storage.

Preferred skills

Observability, monitoring, alerting, capacity planning, performance engineering, MLOps, SaaS and on-prem deployment experience.

Technologies

Kubernetes, AWS, GCP, Terraform, Python, Go

Responsibilities

Define technical roadmap for platform reliability and scalability; lead architecture of deployment platforms; establish reliability engineering standards (SLOs, capacity planning); drive automation to reduce operational toil; lead complex production incident response; mentor engineers; partner with engineering leadership on platform architecture.

Seniority

Staff, hands-on IC with strategic influence and mentorship

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.