Site Reliability Engineer
Core
Operate and scale mission-critical cloud infrastructure, big data workflows, and ML platform components for network assurance analytics.
Role type
Senior Site Reliability Engineer (Data & ML Infrastructure)
Builds
Scalable data pipelines, ML workloads, and cloud-native services on AWS
Domain
Cloud Infrastructure / Big Data / Machine Learning
Deliverable
production ML models | infrastructure
Required skills
AWS operations, Apache Airflow, AWS EMR, Spark, Hadoop, Amazon EKS, Terraform, Python, Linux, distributed systems, observability, incident management, FinOps
Preferred skills
Large-scale SaaS platforms, ML infrastructure optimization, AWS cost management tools, CI/CD systems, configuration management tools
Technologies
AWS, Apache Airflow, AWS EMR, Spark, Hadoop, Amazon EKS, Terraform, Python, CloudWatch, Prometheus, Grafana, Splunk, OpenSearch, Datadog
Responsibilities
Own and operate scalable, reliable, and cost-efficient infrastructure; Manage and optimize large-scale data workflows; Operate and improve Amazon EKS environments; Build and maintain infrastructure automation; Develop Python-based automation and tooling; Drive FinOps practices; Partner with cross-functional teams to improve reliability and performance; Identify infrastructure bottlenecks and cost optimization opportunities; Improve observability and incident response; Support capacity planning and workload efficiency; Lead technical discussions and provide senior-level mentorship.
Seniority
Senior, hands-on IC with leadership responsibilities (via careerplan.io/jobs/2027465-site-reliability-engineer-at-cisco)
