Software Engineer, Platform Reliability Engineering, AiDP
Core
Design, build, and operate mission-critical distributed systems and infrastructure for Apple's GenAI, ML, and big data platforms.
Role type
Senior IC Platform Reliability Engineer (SRE)
Builds
Resilient hybrid cloud infrastructure for inference, data processing, and ML workloads
Domain
Enterprise AI/ML infrastructure on hybrid cloud
Deliverable
infrastructure
Required skills
Distributed systems architecture, Systems programming (Python/Go/Java), Cloud platforms, Containerization, Observability, Incident response, Performance optimization
Preferred skills
SRE/DevOps experience, Big data technologies (Spark/Flink/Iceberg), ML platforms (Ray/MLflow), Linux administration, Security principles
Technologies
Kubernetes, Spark, Flink, Ray, Trino, Iceberg, MLflow
Responsibilities
Design and maintain scalable multi-tenant systems, Own full lifecycle of infrastructure projects, Operate high-throughput mission-critical services, Respond to production incidents and drive post-incident improvements, Lead cross-functional collaboration on platform requirements, Identify operational bottlenecks and implement preventive measures, Establish observability practices and refine operational excellence standards
