Staff Engineer - SRE, Retail and Pharmacy
Core
Architecting and building the foundational infrastructure for reliability, observability, and production readiness across a massive fleet of retail and pharmacy edge nodes and hybrid cloud environments.
Role type
Staff Software Engineer — Site Reliability Engineering (SRE)
Builds
Observability platform, production readiness framework, incident management operating model, chaos engineering program, and developer reliability platform.
Domain
Retail and Pharmacy technology, distributed systems, hybrid cloud, edge computing.
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Distributed systems architecture, observability strategy design, SLO/SLI definition, chaos engineering, fleet-scale deployment management, CI/CD automation, anomaly detection system design, organizational change management.
Preferred skills
AI/LLM integration for alerting, eBPF tooling, OpenTelemetry standards, OPA/Conftest policy enforcement, LitmusChaos.
Technologies
OpenTelemetry, eBPF, Vault, cert-manager, OPA, Conftest, GitHub Actions, LitmusChaos, Chaos Toolkit.
Responsibilities
Architect end-to-end observability platform strategy including hot/warm/cold tier data flows; Define org-wide composite reliability signal frameworks and drive SLI/SLO standardization; Design and own AI-enabled detection pipelines with LLM-based alert summarization; Own the Production Readiness Review (PRR) framework and automated policy enforcement; Design and run organizational chaos engineering and GameDay programs; Lead fleet-scale deployment reliability for thousands of edge nodes.
Seniority
Staff, strategic program ownership and organizational influence