CareerPlanGet AI match score →

Staff Site Reliability Operations Engineer

Bangalore🌐 Remote💼 Full-time🗓 2026-07-08 → 2026-07-31

Core

Lead global platform reliability and next-generation observability strategy on Google Cloud Platform, building intelligent, self-healing infrastructure for Communication Service Providers.

Role type

Staff Site Reliability Operations Engineer (AIOps, GCP, Networking)

Builds

Global platform reliability, observability stack, and self-healing infrastructure for CSPs

Domain

Cloud infrastructure, networking (L1-L7), and AIOps for Communication Service Providers

Deliverable

production ML models | infrastructure

Required skills

Google Kubernetes Engine (GKE) internals, Apache Kafka high-throughput management, PostgreSQL/AlloyDB/BigQuery scaling, Grafana Labs stack (Mimir/Loki/Tempo/Beyla), AIOps anomaly detection, HashiCorp Terraform, Go/Python programming, OSI Layer 1-7 networking protocols (BGP, OSPF, TCP, QUIC, HTTP/3, gRPC)

Preferred skills

Google Cloud architectural best practices (SDN, Cloud Armor, IAM), Linux internals, eBPF-based monitoring, packet analysis tools (Wireshark, tcpdump)

Technologies

Google Cloud Platform, GKE, Kafka, PostgreSQL, AlloyDB, BigQuery, Grafana, Mimir, Loki, Tempo, Beyla, Terraform, Go, Python

Responsibilities

Architect and optimize full-stack network infrastructure (L1-L7), design and scale unified observability platforms, deploy ML models for automated anomaly detection, drive GKE platform engineering, maintain high-throughput Kafka clusters, ensure large-scale database performance and disaster recovery, integrate AIOps for automated incident response, mentor engineers on distributed systems and debugging

Seniority

Staff, hands-on IC with technical leadership and mentorship

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗