Staff Site Reliability Engineer
Core
Architect resilient, self-healing systems and lead multi-region resilience for a high-traffic food delivery platform serving students across the US.
Role type
Staff Site Reliability Engineer
Builds
Production services, multi-region active-standby architecture, and scalable cloud infrastructure for a growing food delivery platform.
Domain
Cloud Infrastructure / SRE / Food Tech
Required skills
Infrastructure as Code (Terraform, Terraspace), Kubernetes (EKS, Helm, KEDA), AWS multi-region architecture, CI/CD (Jenkins, GitHub Actions), Python/Go, distributed monitoring (SLOs, tracing), MySQL/MongoDB/Redis, microservice architecture, PCI compliance.
Preferred skills
Experience with highly trafficked web services, ability to set technical direction across teams, experience with seasonal traffic scaling.
Technologies
AWS, Terraform, Helm, Kubernetes, Jenkins, Python, Go, MySQL, MongoDB, Redis, RabbitMQ, KEDA, Linux
Responsibilities
Architect resilient, self-healing systems and co-own the design of critical production services. Lead multi-region resilience, including active-standby architecture, regional failover readiness, and data replication. Own AWS infrastructure as code using Terraform or Terraspace, Helm, and Helmfile, from design through rollout. Manage the Kubernetes platform, including EKS lifecycle, controller and add-on upgrades, ingress or gateway, and autoscaling. Own observability end to end across logging, metrics, tracing, and alerting pipelines. Develop scaling and capacity strategies for seasonal traffic peaks. Manage platform cloud costs through right-sizing and reservations. Shape incident management by leading responses, postmortems, and architecture reviews.
Seniority
Staff, hands-on IC with technical direction