Site Reliability Engineer, Dublin
Core
Build and enhance massive clusters hosting Virtual Machines and Containers to power Apple's global services (iCloud, Apple Music, etc.) with constant uptime and seamless scaling.
Role type
Senior IC Site Reliability Engineer (Compute Infrastructure)
Builds
Production-grade compute infrastructure (VMs, containers, orchestration) for Apple Services
Domain
Cloud Infrastructure / Virtualization / Containerization
Required skills
Go, Java, Infrastructure-as-a-Service (IaaS), Cloud Operations, Virtualization, Containerization, CI/CD, Observability, Incident Response, Capacity Planning
Preferred skills
OpenStack, CloudStack, Linux System Virtualization (Libvirt, QEMU, KVM), Telemetry Implementation, Distributed Systems
Technologies
Go, Java, OpenStack, CloudStack, Libvirt, QEMU, KVM, Splunk, Grafana, Prometheus
Responsibilities
Design and develop tooling, frameworks, and automation in Go and Java for compute infrastructure; Define and implement SLOs/SLIs and build observability pipelines; Lead incident response, triage, and root cause analysis; Develop and maintain infrastructure-as-code and CI/CD pipelines; Contribute to compute platform architecture through design reviews and capacity planning; Partner cross-functionally to embed reliability into the development lifecycle.