Site Reliability Engineer, Enterprise Technology Services
Core
Design, build, and operate large-scale Identity Management Platform services ensuring ultra-high availability for authentication, authorization, and user provisioning across Apple's ecosystem.
Role type
Senior Site Reliability Engineer (Identity Management)
Builds
Distributed identity services, automation frameworks, observability stacks, and resilience platforms for critical authentication and authorization transactions.
Domain
Cloud/On-premise Infrastructure, Identity & Access Management, Distributed Systems
Required skills
Distributed systems architecture, SRE principles (SLIs/SLOs/SLAs), observability stack design, automation engineering, incident management, security compliance, capacity planning, disaster recovery, fraud prevention, ML/GenAI for anomaly detection, CI/CD pipelines, Infrastructure as Code, multi-database management, event-driven architectures.
Preferred skills
Chaos engineering, canary releases, cryptography, OAuth/SAML/SSO, governance frameworks, ML/GenAI for operational efficiency.
Technologies
Python, Java, Go, Bash, Ansible, Prometheus, Grafana, Datadog, OpenTelemetry, ELK, Kafka, RabbitMQ, Helm, CRD, Git.
Responsibilities
Define and implement SLIs/SLOs/SLAs for platform reliability; design and manage resilient distributed systems across cloud and on-premise; lead incident response and post-mortems; develop large-scale automation and tooling; ensure security posture and compliance with industry standards; partner with engineering teams to optimize system performance and architecture.
Seniority
Senior, hands-on IC
