Staff Site Reliability Operations Engineer
Core
Lead global platform reliability and next-generation observability strategy on Google Cloud Platform, building intelligent, self-healing infrastructure for Communication Service Providers.
Role type
Staff Site Reliability Operations Engineer (AIOps, GCP, Networking)
Builds
Global platform reliability, observability stack, and self-healing infrastructure for CSPs
Domain
Cloud infrastructure, networking (L1-L7), and AIOps for Communication Service Providers
Deliverable
production ML models | infrastructure
Required skills
Google Kubernetes Engine (GKE) internals, Apache Kafka high-throughput management, PostgreSQL/AlloyDB/BigQuery scaling, Grafana Labs stack (Mimir/Loki/Tempo/Beyla), AIOps anomaly detection, HashiCorp Terraform, Go/Python programming, OSI Layer 1-7 networking protocols (BGP, OSPF, TCP, QUIC, HTTP/3, gRPC)
Preferred skills
Google Cloud architectural best practices (SDN, Cloud Armor, IAM), Linux internals, eBPF-based monitoring, packet analysis tools (Wireshark, tcpdump)
Technologies
Google Cloud Platform, GKE, Kafka, PostgreSQL, AlloyDB, BigQuery, Grafana, Mimir, Loki, Tempo, Beyla, Terraform, Go, Python
Responsibilities
Architect and optimize full-stack network infrastructure (L1-L7), design and scale unified observability platforms, deploy ML models for automated anomaly detection, drive GKE platform engineering, maintain high-throughput Kafka clusters, ensure large-scale database performance and disaster recovery, integrate AIOps for automated incident response, mentor engineers on distributed systems and debugging
Seniority
Staff, hands-on IC with technical leadership and mentorship