Staff Site Reliability Engineer (SRE) (Hybrid)
Core
Provide technical leadership for the reliability, scalability, and operational architecture of Splunk Agent Observability's platform to ensure AI agents behave as intended.
Role type
Staff Site Reliability Engineer (SRE)
Builds
Scalable, cost-effective evaluation and guardrails for AI agents; deployment platforms for cloud and air-gapped environments.
Domain
AI resilience, observability, cloud infrastructure, and distributed systems.
Deliverable
production ML models | infrastructure
Required skills
Site Reliability Engineering, Kubernetes, distributed systems design, cloud platforms (AWS/GCP), CI/CD automation, incident response, technical leadership, Python, Go, Infrastructure as Code (Terraform), networking, databases, storage.
Preferred skills
Observability, monitoring, alerting, capacity planning, performance engineering, MLOps, SaaS and on-prem deployment experience.
Technologies
Kubernetes, AWS, GCP, Terraform, Python, Go
Responsibilities
Define technical roadmap for platform reliability and scalability; lead architecture of deployment platforms; establish reliability engineering standards (SLOs, capacity planning); drive automation to reduce operational toil; lead complex production incident response; mentor engineers; partner with engineering leadership on platform architecture.
Seniority
Staff, hands-on IC with strategic influence and mentorship
