Senior Site Reliability Engineer
Core
Own the end-to-end AWS infrastructure platform for IoT controllers and gateways serving the food and beverage industry, ensuring reliability, scale, and rapid incident response.
Role type
Senior Site Reliability Engineer (Infrastructure & Platform)
Builds
AWS infrastructure (ECS/EKS), microservices, distributed systems, CI/CD pipelines, and production databases
Domain
IoT, Food & Beverage, Cloud Infrastructure
Required skills
AWS architecture, Linux, networking, container orchestration (ECS/EKS), infrastructure-as-code, CI/CD, observability, high-availability database design, incident management, technical mentorship
Preferred skills
AI tools and agents for incident diagnosis and runbook automation
Technologies
AWS (ECS, EKS, CDK, CloudFormation, CloudWatch), ArgoCD, Prometheus, Grafana
Responsibilities
Lead design and delivery of AWS infrastructure for containers and microservices; Own reliability workstreams from architecture to operation; Define monitoring, alerting, and automation standards; Set CI/CD and release strategies (blue/green, canary); Ensure production database reliability (HA, backup, recovery); Lead major incident response and root cause analysis; Mentor junior and intermediate engineers
Seniority
Senior, hands-on IC with mentorship responsibilities