Staff Site Reliability Operations Engineer
Core
Lead global platform reliability and next-generation observability strategy for a cloud-first AI platform serving Communication Service Providers.
Role type
Staff Site Reliability Operations Engineer (SRE)
Builds
Intelligent, self-healing infrastructure and unified observability platform on Google Cloud Platform
Domain
Cloud Infrastructure, Networking, Observability, AIOps
Deliverable
production ML models | infrastructure
Required skills
Google Kubernetes Engine (GKE), Apache Kafka, PostgreSQL, AlloyDB, BigQuery, Grafana stack (Mimir, Loki, Tempo, Beyla), AIOps/ML anomaly detection, HashiCorp Terraform, Go, Python, OSI Layer 1-7 networking, GitOps
Preferred skills
Google Cloud Platform (GCP) architectural best practices, Linux internals, eBPF-based monitoring, packet analysis tools
Technologies
Google Cloud Platform, GKE, Kafka, PostgreSQL, AlloyDB, BigQuery, Grafana, Mimir, Loki, Tempo, Beyla, Terraform, Go, Python
Responsibilities
Architect and optimize full-stack network infrastructure (L1-L7); Design and scale unified observability platform using Grafana stack; Deploy machine learning models for automated anomaly detection and alerting; Drive architecture and scaling of production GKE clusters; Tune and maintain high-throughput Apache Kafka clusters; Ensure performance and disaster recovery of large-scale data tiers; Integrate AIOps insights for automated incident response; Mentor engineers on distributed systems and advanced debugging
Seniority
Staff, technical leadership & mentorship