SRE Manager
Core
Lead and manage Site Reliability Engineering (SRE) operations for a global SaaS platform, ensuring high availability, resilience, and automated infrastructure for order fulfillment systems.
Role type
Senior SRE Manager (Technical Leadership)
Builds
Scalable SaaS platforms for foodservice, retail, and wholesale distribution enabling contactless order pickup.
Domain
SaaS / Cloud Infrastructure / IoT / Big Data
Deliverable
production ML models | product features | infrastructure
Required skills
SRE team leadership, SLO/SLI definition, blameless postmortem culture, Infrastructure as Code (IaC), CI/CD pipeline design, microservice architecture, system security, network traffic management, monitoring and alerting, APM tooling, automation scripting.
Preferred skills
Experience with reverse proxies (Nginx, Caddy, Traefik), capacity planning and reporting, public code repository maintenance.
Technologies
AWS, Azure, GCP, Kubernetes, Terraform, Ansible, Jenkins, GitLab CI/CD, GitHub Actions, Prometheus, Datadog, Dynatrace, AppDynamics, New Relic, Node.js, Python, Java, Go, C++, PHP, Perl, Unix shell, RabbitMQ, Kafka, ActiveMQ, Docker, Nginx, Angular, Blockchain, IoT, Big Data.
Responsibilities
Oversee day-to-day SRE team operations; Define and refine SLOs/SLIs to exceed business SLAs; Foster blameless postmortem culture; Design and maintain build/release infrastructure and CI/CD pipelines; Solve complex operational problems via software and automation; Collaborate on system design and platform resilience; Analyze capacity and performance for senior leadership; Administer, monitor, and secure 24x7 SaaS environments.
Seniority
Senior, hands-on IC with direct reports