CareerPlanGet AI match score →

Lead Site Reliability Engineer, Network Assurance Data Platform - Cisco ThousandEyes

Bangalore, India💼 Full-time🗓 2026-06-02 → 2026-07-31

Core

Lead SRE for the Network Assurance Data Platform, ensuring reliability, scalability, and security of cloud and big data platforms supporting ML/AI initiatives.

Role type

Senior IC Site Reliability Engineering (SRE) Technical Leader

Builds

Cloud and big data infrastructure for ML/AI workloads, SaaS systems operating at multi-region scale

Domain

Cloud infrastructure, Big Data, Machine Learning/AI

Deliverable

production ML models | infrastructure

Required skills

Cloud architecture (AWS), Infrastructure as Code (Terraform, Kubernetes/EKS), Big Data orchestration (Hadoop ecosystem, Spark, Airflow, EMR, SageMaker), Python/Go programming, Root cause analysis, Technical leadership, Automation

Preferred skills

Unix/Linux kernel and system libraries, Observability tools (Prometheus, Grafana, Thanos, CloudWatch, OpenTelemetry, ELK), Software architecture at scale

Technologies

AWS, Terraform, Kubernetes, EKS, Spark, Hive, HDFS, Gobblin, Airflow, EMR, SageMaker, Python, Go, Prometheus, Grafana, Thanos, CloudWatch, OpenTelemetry, ELK

Responsibilities

Design and optimize cloud/data infrastructure for high availability and scalability; Collaborate with cross-functional teams to create secure, scalable solutions; Troubleshoot production issues and perform root cause analyses; Lead architectural vision and technical strategy; Mentor teams and foster engineering excellence; Engage with stakeholders to translate use cases into actionable insights; Develop strategic roadmaps and processes for enterprise-scale deployment

Seniority

Senior, hands-on IC with leadership responsibilities

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗