Lead Site Reliability Engineer, Network Assurance Data Platform - Cisco ThousandEyes
Core
Lead SRE for the Network Assurance Data Platform, ensuring reliability, scalability, and security of cloud and big data platforms supporting ML/AI initiatives.
Role type
Senior IC Site Reliability Engineering (SRE) Technical Leader
Builds
Cloud and big data infrastructure for ML/AI workloads, SaaS systems operating at multi-region scale
Domain
Cloud infrastructure, Big Data, Machine Learning/AI
Deliverable
production ML models | infrastructure
Required skills
Cloud architecture (AWS), Infrastructure as Code (Terraform, Kubernetes/EKS), Big Data orchestration (Hadoop ecosystem, Spark, Airflow, EMR, SageMaker), Python/Go programming, Root cause analysis, Technical leadership, Automation
Preferred skills
Unix/Linux kernel and system libraries, Observability tools (Prometheus, Grafana, Thanos, CloudWatch, OpenTelemetry, ELK), Software architecture at scale
Technologies
AWS, Terraform, Kubernetes, EKS, Spark, Hive, HDFS, Gobblin, Airflow, EMR, SageMaker, Python, Go, Prometheus, Grafana, Thanos, CloudWatch, OpenTelemetry, ELK
Responsibilities
Design and optimize cloud/data infrastructure for high availability and scalability; Collaborate with cross-functional teams to create secure, scalable solutions; Troubleshoot production issues and perform root cause analyses; Lead architectural vision and technical strategy; Mentor teams and foster engineering excellence; Engage with stakeholders to translate use cases into actionable insights; Develop strategic roadmaps and processes for enterprise-scale deployment
Seniority
Senior, hands-on IC with leadership responsibilities