Site Reliability Engineer
Core
Keep data and AI platforms running reliably, efficiently, and securely to power global retail decisions.
Role type
Site Reliability Engineer (SRE)
Builds
Data and AI platforms on Google Cloud Platform (GCP) and Azure
Domain
Retail / Cloud Infrastructure / Data Engineering
Deliverable
production ML models | infrastructure
Required skills
GCP services (GKE, Cloud Run, BigQuery, Pub/Sub, GCS), Infrastructure-as-Code (Terraform, Helm), observability tooling (metrics, logging, tracing, alerting), SLO/SLI concepts, scripting (Python, Bash, Go), container orchestration (Kubernetes), CI/CD pipelines
Preferred skills
Retail/e-commerce environment experience, Google SRE principles, AI/ML platform operations, FinOps
Technologies
GCP, Azure, Kubernetes, GKE, Terraform, Helm, Cloud Monitoring, Datadog, Prometheus, Grafana, ArgoCD, GitHub Actions
Responsibilities
Monitor production systems using observability tooling to detect and triage issues; participate in on-call rotations and respond to incidents; build automation to reduce operational toil; maintain SLO dashboards and alerting thresholds; operate and maintain workloads on GCP and Azure; apply Infrastructure-as-Code practices to manage infrastructure changes.
Seniority
Mid-Senior, hands-on IC
