Site Reliability Engineer
Core
Build, maintain, and improve automation solutions and production platform services for Cisco's internal network using DevOps and SRE practices.
Role type
Senior Site Reliability Engineer (Network Services)
Builds
Internal network services, automation solutions, and platform infrastructure for Cisco's own environment.
Domain
Enterprise Networking & Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Python/Shell scripting, Linux administration (RHEL), Container orchestration (Kubernetes/Podman), CI/CD pipelines, API integration, Observability (monitoring/logging/alerting), Incident response, Root cause analysis, SLI/SLO definition
Preferred skills
Public cloud platforms, Cisco platforms (Catalyst Center, NSO), Disaster recovery planning, Enterprise networking (TCP/IP, routing, switching), ServiceNow/Jira, OpenTelemetry/AIOps, AI-assisted coding tools
Technologies
Python, Shell, Linux, RHEL, Podman, Kubernetes, Git, Jenkins, GitHub Copilot, Docker, ServiceNow, Jira, OpenTelemetry
Responsibilities
Develop and maintain automation solutions; Support production platform services for reliability and performance; Create runbooks and disaster recovery processes; Lead incident response and postmortems; Improve observability and operational metrics; Partner with cross-functional teams to identify failure points.
Seniority
Mid-Senior, hands-on IC