Senior Site Reliability Engineer
Core
Operate and scale the core data infrastructure for Cisco ThousandEyes' Network Assurance Data Platform, supporting large-scale data pipelines, ML workloads, and cloud-native services on AWS.
Role type
Senior Site Reliability Engineer (Data Infrastructure)
Builds
Scalable, reliable, and cost-efficient data infrastructure and ML platforms on AWS
Domain
Cloud Infrastructure / Big Data / Machine Learning
Deliverable
production ML models | infrastructure
Required skills
AWS operations, Apache Airflow, AWS EMR, Spark, Hadoop, Amazon EKS, Kubernetes, Terraform, Python, Linux, distributed systems, observability, incident management, FinOps
Preferred skills
Large-scale SaaS platforms, ML infrastructure optimization, AWS cost management tools, CI/CD systems, configuration management tools, FinOps certification
Technologies
AWS, Apache Airflow, AWS EMR, Spark, Hadoop, Amazon EKS, Terraform, Python, CloudWatch, Prometheus, Grafana, Splunk, OpenSearch, Datadog
Responsibilities
Own and operate scalable infrastructure for the Network Assurance Data Platform; Manage and optimize large-scale data workflows; Operate Amazon EKS environments for containerized and ML workloads; Build infrastructure automation using Terraform; Develop Python-based automation and operational tooling; Drive FinOps practices for cost visibility and optimization; Partner with cross-functional teams to improve reliability and performance; Identify infrastructure bottlenecks and cost optimization opportunities; Improve observability and incident response; Lead technical discussions and mentor engineers.
Seniority
Senior, hands-on IC with technical leadership