CareerPlanSign in

Technical Operations & Site Reliability Engineer, Customer Systems

Sunnyvale, United States of America💼 Full-time🗓 2026-03-30 → 2026-09-28

Core

Design and maintain automation solutions for large-scale, globally distributed customer systems to ensure reliability, availability, and performance.

Role type

Senior IC Technical Operations & Site Reliability Engineer

Builds

Automated monitoring, incident response, and operational workflows for business-critical global applications

Domain

Technology / Distributed Systems / Customer Experience

Deliverable

production ML models

Required skills

Java/JEE, REST, Swift/Objective C, Python, Go, Bash, Linux kernel management, networking protocols (HTTP, DNS, TCP/IP, ICMP), log analysis, incident response, AI/LLM model training and optimization, database schema design

Preferred skills

Microservices architecture, messaging brokers, versioning strategies, 24x7 global operations experience, system health monitoring strategy

Technologies

Hubble, ExtraHop, Splunk, Linux, AI & LLM models

Responsibilities

Manage large-scale production outages and lead incident response; Develop tools to automate repetitive operational tasks; Plan and execute system health monitoring and communication; Partner with teams to improve reliability and processes; Create and maintain technical documentation and training materials

Seniority

Senior, hands-on IC

Sourced via apple · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.