CareerPlanGet AI match score →

Staff Site Reliability Operations Engineer

Bangalore💼 Full-time🗓 2026-07-21 → 2026-07-31

Core

Lead global platform reliability and next-generation observability strategy for a cloud-first AI platform serving Communication Service Providers.

Role type

Staff Site Reliability Operations Engineer (SRE)

Builds

Intelligent, self-healing infrastructure and unified observability platform on Google Cloud Platform

Domain

Cloud Infrastructure, Networking, Observability, AIOps

Deliverable

production ML models | infrastructure

Required skills

Google Kubernetes Engine (GKE), Apache Kafka, PostgreSQL, AlloyDB, BigQuery, Grafana stack (Mimir, Loki, Tempo, Beyla), AIOps/ML anomaly detection, HashiCorp Terraform, Go, Python, OSI Layer 1-7 networking, GitOps

Preferred skills

Google Cloud Platform (GCP) architectural best practices, Linux internals, eBPF-based monitoring, packet analysis tools

Technologies

Google Cloud Platform, GKE, Kafka, PostgreSQL, AlloyDB, BigQuery, Grafana, Mimir, Loki, Tempo, Beyla, Terraform, Go, Python

Responsibilities

Architect and optimize full-stack network infrastructure (L1-L7); Design and scale unified observability platform using Grafana stack; Deploy machine learning models for automated anomaly detection and alerting; Drive architecture and scaling of production GKE clusters; Tune and maintain high-throughput Apache Kafka clusters; Ensure performance and disaster recovery of large-scale data tiers; Integrate AIOps insights for automated incident response; Mentor engineers on distributed systems and advanced debugging

Seniority

Staff, technical leadership & mentorship

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗