Staff Site Reliability Engineer
Core
Design, build, and operate highly scalable, reliable, and secure multi-cloud infrastructure powering production systems and enabling microservice architectures.
Role type
Staff Site Reliability Engineer (Infrastructure & Cloud)
Builds
Production systems, container platforms (EKS/GKE), and microservice-based applications across AWS and GCP.
Domain
Cloud Infrastructure, Multi-cloud (AWS/GCP), Kubernetes, Security
Required skills
Kubernetes (EKS/GKE), Cloud Infrastructure (AWS/GCP), Infrastructure as Code (Terraform/Ansible), Python/Go/Shell scripting, Observability (Prometheus/Grafana/ELK), CI/CD (GitOps), Database/Caching management (Redis/PostgreSQL/MySQL), Container Security
Preferred skills
ECS to EKS/GKE migration experience, SLO/SLI definition, Blameless postmortems, Cross-team project leadership, SaaS/High-scale environment experience
Technologies
AWS, GCP, Kubernetes, EKS, GKE, Terraform, Ansible, Python, Go, Shell, Prometheus, Grafana, ELK, Loki, OpenTelemetry, ArgoCD, GitLab CI, Spinnaker, Redis, RDS, Cloud SQL, PostgreSQL, MySQL, HashiCorp Vault, AWS Secrets Manager, Google Secret Manager
Responsibilities
Design and operate scalable multi-cloud infrastructure; Lead reliability and modernization initiatives (e.g., container platform migrations); Serve as technical authority in Kubernetes and cloud infrastructure; Partner with dev teams to enable microservices; Implement IaC for automation; Drive observability and cost efficiency improvements; Champion SRE best practices (SLOs/SLIs, postmortems); Lead complex technical projects from conception to completion; Mentor engineers; Participate in on-call rotation.
Seniority
Staff, hands-on IC with significant mentorship and strategic leadership (via careerplan.io/jobs/7840544-staff-site-reliability-engineer-at-okta)