Sr Staff Site Reliability Engineer
Core
Operate and maintain large-scale, global production environments across multiple cloud providers (GCP, AWS, Azure) for a cybersecurity company serving tens of thousands of enterprise customers.
Role type
Sr Staff Site Reliability Engineer (IC)
Builds
High-availability, multi-cloud infrastructure and observability systems
Domain
Cybersecurity / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Terraform, Python, Prometheus, Grafana, CI/CD, GitOps, incident response, distributed system troubleshooting
Preferred skills
Multi-cloud architecture design, automation tooling development, asynchronous communication standards
Technologies
Kubernetes, Terraform, GCP, AWS, Azure, Prometheus, Grafana, PagerDuty, GitLab CI, GitHub Actions, Jenkins, Flux
Responsibilities
Own and operate large-scale global production environments; monitor and resolve incidents via automated alerting; design and improve monitoring/observability systems; develop automation and tooling in Python; collaborate with internal teams on system reliability; manage on-call rotation for daytime hours and occasional weekends.
Seniority
Sr Staff, hands-on IC (via careerplan.io/jobs/JR-018703-sr-staff-site-reliability-engineer-at-paloaltonetworks)