Site Reliability Engineer
Core
Implement and maintain DevOps practices, manage the lifecycle of machine learning models in production, and ensure scalability, reliability, and performance of Apple B2B systems.
Role type
Senior IC Site Reliability Engineer (SRE)
Builds
Apple B2B systems, machine learning model pipelines, and container orchestrating systems
Domain
Technology / Cloud Infrastructure / Machine Learning Operations
Deliverable
production ML models
Required skills
Java, Python, Kubernetes, EKS, Splunk, Grafana, Prometheus, Oracle, MongoDB, JVM, Operating Systems, Network Protocols, HTTP/S, TCP, DNS, SSL/TLS, SSH/SFTP, PKI, X509, PGP
Preferred skills
WebMethods Integration server, middleware platforms
Technologies
Kubernetes, EKS, Splunk, Grafana, Prometheus, Oracle, MongoDB, Java, Python
Responsibilities
Manage the lifecycle of machine learning models in production and non-production environments; implement and coordinate telemetry using monitoring and observability tools; solve and resolve issues in Kubernetes from both OS and application perspectives; perform performance tuning of applications and databases.
Seniority
Senior, hands-on IC
