CareerPlanSign in

Site Reliability Engineer (Edge Services), Infrastructure Services

Sunnyvale, United States of America💼 Full-time🗓 2026-05-18 → 2026-09-28

Core

Design and implement next-generation observability, alerting, and self-healing systems for production ecosystems to ensure resilience and scalability.

Role type

Senior Site Reliability Engineer (Edge Services)

Builds

High-cardinality observability data, automated reliability frameworks, and self-healing distributed systems

Domain

Cloud Infrastructure, Edge Services, Distributed Systems

Required skills

Linux internals, HTTP/2/3 (QUIC), HTTPS/TLS, Python, Go, Prometheus, Grafana, ClickHouse, Data Structures and Algorithms, SLIs/SLOs/Error Budgets

Preferred skills

Terraform, Ansible, Pulumi, Kubernetes, blameless post-mortems, Generative AI tools for observability/debugging

Technologies

Prometheus, Grafana, ClickHouse, Python, Go, Terraform, Ansible, Pulumi, Kubernetes, AWS, GCP, Azure

Responsibilities

Design observability and alerting strategies prioritizing high-cardinality data; build self-healing systems and reduce toil via automation; partner with dev teams to integrate reliability into CI/CD; debug protocol-level issues and optimize traffic flow; manage cloud environments and containerized workloads; lead post-mortems to harden systems against failures.

Seniority

Senior, hands-on IC

Sourced via apple · Listed on CareerPlan, which tracks 848,000+ jobs from 20+ sources.