Staff Site Reliability Engineer
Core
Own availability and operational excellence for planet-scale observability and security products by optimizing operations, cloud resource usage, and developer velocity.
Role type
Staff Site Reliability Engineer (Product Area Focus)
Builds
Observability and security products for digital teams
Domain
Cloud-native security and observability
Deliverable
production ML models | product features | infrastructure
Required skills
Cloud native application development, debugging and troubleshooting, AWS networking/compute/storage/managed services, CI/CD tooling (Kubernetes, Terraform, Ansible, Jenkins), full lifecycle service support, Infrastructure as Code, production code authoring (Java/Scala/Go), Linux systems, cloud-native security, agile frameworks
Preferred skills
Planet scale product development, streaming technologies (Kafka, Kafka Streams, KSQL), JVM workload tuning
Technologies
AWS, Kubernetes, Terraform, Ansible, Jenkins, Kafka, Kafka Streams, KSQL, Java, Scala, Go, Python
Responsibilities
Maintain and execute a reliability roadmap for the product area, define and manage SLOs for multiple teams, participate in on-call rotations to improve operational experience, write code and automation to reduce toil and improve security, facilitate blame-free root cause analysis, hire and mentor new team members
Seniority
Staff, hands-on IC with mentorship responsibilities