Site Reliability Engineer II, APAC
Core
Balance development velocity with reliability for Apple device management services by implementing observability, automating toil, and utilizing agentic AI tools.
Role type
Senior IC Site Reliability Engineer (SRE)
Builds
Production services for Apple device management (Mac, iPad, iPhone, Apple TV) on AWS
Domain
Cloud Infrastructure (AWS) + Observability + AI Engineering
Deliverable
production ML models | infrastructure
Required skills
Production troubleshooting across stack, AWS operations, Observability tools, Automation scripting, Technical documentation, Agentic AI tool usage, AI output verification
Preferred skills
Infrastructure as Code, CI/CD tooling, Shared AI asset contribution
Technologies
AWS (EC2, S3, EKS, RDS/Aurora, CloudFront), Grafana, Prometheus, LogicMonitor, Python/Go/Java, Terraform, GitHub Actions, Jenkins, Claude Code, Cursor, Copilot
Responsibilities
Implement and maintain service level objectives and error budgets; Investigate production issues using AI to correlate logs, metrics, and code; Produce technical documentation, runbooks, and postmortems; Identify and eliminate toil through automation and AI agents; Refine conditions for AI agents (task definitions, guardrails, MCP servers); Participate in on-call rotation for production incidents; Informally mentor less-experienced engineers on debugging and AI usage.
Seniority
Senior, hands-on IC