Site Reliability Engineer / Devops — Retail Engineering
Core
Drive reliability, deployment, and scalability of compute platforms across on-premises and hybrid cloud environments for Apple's retail and marketing systems.
Role type
Senior Site Reliability Engineer (Infrastructure & Platform)
Builds
Compute platforms, application stacks, and CI/CD pipelines for retail and marketing experiences
Domain
Retail technology, hybrid cloud infrastructure, large-scale distributed systems
Required skills
Site Reliability Engineering, distributed systems design, Infrastructure as Code, container orchestration, CI/CD pipeline optimization, observability stack design, troubleshooting complex systems, scripting/programming (Python, Java, Go, Bash, Ansible)
Preferred skills
SRE principles (error budgeting, SLO/SLI/SLA), advanced Java/Python/Go programming, relational/NoSQL/OLAP databases, event-driven streaming (Kafka, RabbitMQ), on-call rotation management, root cause analysis, enterprise security standards, cryptography, authentication protocols (OAuth, SAML, SSO)
Technologies
Kubernetes, Docker, Terraform, Ansible, Prometheus, Grafana, Datadog, OpenTelemetry, ELK, Kafka, RabbitMQ, Java, Python, Go, Bash
Responsibilities
Deploy, support, and monitor compute platforms and application stacks; build Infrastructure as Code; optimize container orchestration; streamline CI/CD delivery pipelines; champion automation and operational excellence; develop and fix applications on failures
Seniority
Senior, hands-on IC