Senior Site Reliability Engineer (SRE) (Hybrid)
Core
Build, operate, and improve the reliability, scalability, and operational excellence of Splunk Agent Observability's deployment platform and production infrastructure for cloud and air-gapped customer deployments.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Scalable, cost-effective evaluation and guardrails for AI agent resilience; deployment platforms and production infrastructure.
Domain
AI/ML Observability, Cloud Infrastructure, Kubernetes
Deliverable
production ML models | infrastructure
Required skills
Kubernetes production operations, CI/CD platform development, cloud platform experience (AWS/GCP), Python/Go programming, Infrastructure as Code (Terraform), database tuning, distributed systems debugging
Preferred skills
MLOps experience, observability/logging/alerting platforms, networking fundamentals (VPCs, DNS, routing, load balancing), air-gapped/on-prem environment deployment
Technologies
Kubernetes, Helm, Terraform, Python, Go, AWS, GCP
Responsibilities
Operate and improve Kubernetes-based production infrastructure; own customer deployments across cloud and air-gapped environments; build and improve deployment observability, monitoring, logging, and alerting; participate in production incident response and root cause analysis; design and develop internal tooling; debug complex production issues spanning Kubernetes, networking, storage, and application layers; collaborate with software engineers and customers to design secure, scalable, and reliable deployment architectures
Seniority
Senior, hands-on IC
