Site Reliability Engineer, Enterprise Technology Services
Core
Operate and improve a multi-petabyte, highly-available Big Data ecosystem supporting Apple's global manufacturing operations and product launches.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Exabyte-scale data infrastructure for manufacturing, identity, and device security platforms
Domain
Manufacturing operations / Big Data / Cloud Infrastructure
Deliverable
production ML models
Required skills
Linux systems administration, Python/Go/Bash programming, Cloud platforms (AWS/GCP), AI/ML for IT operations, Incident management, Infrastructure-as-code
Preferred skills
Kubernetes, CI/CD pipelines, Monitoring (Grafana/Prometheus), Big data technologies (Kafka/Druid), Relational databases (MySQL/PostgreSQL), Configuration management (Ansible)
Technologies
AWS, GCP, Python, Go, Bash, Kubernetes, Terraform, Prometheus, Kafka, MySQL, PostgreSQL, ArgoCD, Jenkins, Ansible
Responsibilities
Own reliability, performance, and scalability of services; Build AIOps capabilities including ML-driven alerting and anomaly detection; Instrument services for deep observability with dashboards and SLOs; Drive automation to reduce operational toil; Partner with incident management for response and post-incident reviews; Contribute to infrastructure-as-code practices and platform architecture
Seniority
Senior, hands-on IC