Sr Staff Site Reliability Engineer
Core
Operate and improve large-scale, global production environments across multiple cloud providers to prevent cyberattacks and ensure system reliability.
Role type
Sr Staff Site Reliability Engineer (IC)
Builds
Large-scale, multi-cloud production environments supporting tens of thousands of enterprise customers
Domain
Cybersecurity / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Terraform, Python, Prometheus, Grafana, CI/CD, GitOps, incident response, distributed system troubleshooting
Preferred skills
Multi-cloud expertise (GCP, AWS, Azure), automation development, asynchronous communication standards
Technologies
Kubernetes, Terraform, GCP, AWS, Azure, Prometheus, Grafana, PagerDuty, GitLab CI, GitHub Actions, Jenkins, Flux
Responsibilities
Own and operate large-scale global production environments across multiple cloud providers; Monitor, investigate, and resolve incidents triggered by automated alerting systems; Design, deploy, and improve monitoring and observability systems; Collaborate with internal teams to ensure system reliability and performance; Develop and maintain automation and tooling in Python; Champion asynchronous communication and documentation standards; Handle on-call responsibilities including daytime hours and occasional weekends/holidays.
Seniority
Sr Staff, hands-on IC