Senior Site Reliability Engineer (Cortex)
Core
Operate and maintain large-scale, multi-cloud production environments supporting tens of thousands of enterprise customers in a cybersecurity organization.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Global production environments across GCP, AWS, and Azure
Domain
Cybersecurity / Cloud Infrastructure
Deliverable
infrastructure
Required skills
Kubernetes, Terraform, Python, Prometheus, Grafana, CI/CD, GitOps, incident response, distributed systems troubleshooting
Preferred skills
Multi-cloud architecture, automation development, operational excellence
Technologies
Kubernetes, Terraform, GCP, AWS, Azure, Prometheus, Grafana, PagerDuty, GitLab CI, GitHub Actions, Jenkins, Flux
Responsibilities
Own and operate large-scale global production environments; monitor and resolve incidents via automated alerting; design and improve monitoring and observability systems; develop and maintain automation and tooling; collaborate with internal teams on system reliability; manage on-call rotations.
Seniority
Senior, hands-on IC