CareerPlanGet AI match score →

Customer Reliability Engineer, Hypershield (remote)

51 Locations🌐 Remote💼 Full-time🗓 2026-07-10 → 2026-07-31

Required skills

Experience supporting enterprise customers in an escalation capacity, including diagnosing and resolving complex production incidents under SLA pressure in unfamiliar environments, Experience operating and troubleshooting Cisco Nexus / NX-OS in production; equivalent depth on another major vendor accepted, Prior experience to localize failures across a layered data-center architecture spanning switching/forwarding, services/enforcement, and control-plane domains, Linux operations experience at the command line, including production troubleshooting, with working exposure to containers or Kubernetes

Preferred skills

Direct experience using network troubleshooting tooling as a primary diagnostic method, including packet capture and flow-telemetry analysis (NetFlow/IPFIX), Working knowledge of enterprise virtualization, sufficient to troubleshoot a VM-based appliance deployment; vSphere is the current deployment target, Operational Kubernetes and Helm proficiency, including diagnosing failures beyond the workload level: TLS certificates and service-account authentication, API-server connectivity, service exposure, persistent storage for stateful workloads, and custom resources and operators, Working knowledge of VXLAN EVPN fabrics, including the Smart Switch's placement within them and the ability to isolate faults across the fabric, Network segmentation and firewall policy design (zone-based or microsegmentation), and the tradeoffs vs. traditional NGFWs, Familiarity with the NetOps/NetSecOps operating split in data-center security, Experience driving diagnosis and remediation through a customer's own team, working incidents in environments with no direct access where the customer accomplishes the steps, Experience with NX-OS automation and APIs (NX-API, NETCONF/RESTCONF, gNMI, or Ansible); familiarity with Cisco Nexus Dashboard a plus, Demonstrated ability to communicate incident status, root cause, and remediation clearly to both technical and executive audiences, verbally and in writing, CCNP Data Center, CCNP Enterprise, DevNet Professional, CCIE Data Center, CCIE Enterprise, or DevNet Expert

Technologies

Cisco Nexus / NX-OS, Linux, containers, Kubernetes, Helm, VXLAN EVPN fabrics, vSphere, network segmentation, firewall policy design, NetOps/NetSecOps, NX-OS automation and APIs, Cisco Nexus Dashboard

Responsibilities

Own Hypershield cases escalated from Cisco TAC through to resolution, engaging customers directly as the incident requires, Diagnose complex production failures through the Hypershield surface: the N9300 Smart Switch fabric and the on-premises Kubernetes controller that manages its security policy, Localize faults across the layered architecture: switching and forwarding, security services and enforcement, and the control plane, Develop a deep understanding of each customer's architecture and configuration, and diagnose failures in unfamiliar production environments, Reproduce customer failures, partner with engineering to drive fixes, and own the fix back to the customer, Convert individual cases into systemic improvements: runbooks, diagnostics, knowledge-base content, and product feedback to engineering, Help build the team's proactive view of customer health, developing new monitoring, tooling, and reliability practices as the installed base grows

Seniority

Bachelor's + 8 years of experience, Master's + 6 years, or equivalent industry experience

Domain

Networking, security, observability, data-center, Kubernetes, enterprise solutions

Rewrite
## Responsibilities - Own Hypershield cases escalated from Cisco TAC through to resolution, engaging customers directly as the incident requires - Diagnose complex production failures through the Hypershield surface: the N9300 Smart Switch fabric and the on-premises Kubernetes controller that manages its security policy - Localize faults across the layered architecture: switching and forwarding, security services and enforcement, and the control plane - Develop a deep understanding of each customer's architecture and configuration, and diagnose failures in unfamiliar production environments - Reproduce customer failures, partner with engineering to drive fixes, and own the fix back to the customer - Convert individual cases into systemic improvements: runbooks, diagnostics, knowledge-base content, and product feedback to engineering - Help build the team's proactive view of customer health, developing new monitoring, tooling, and reliability practices as the installed base grows ## Requirements - Bachelor's + 8 years of experience, Master's + 6 years, or equivalent industry experience - Experience supporting enterprise customers in an escalation capacity, including diagnosing and resolving complex production incidents under SLA pressure in unfamiliar environments - Experience operating and troubleshooting Cisco Nexus / NX-OS in production; equivalent depth on another major vendor accepted - Prior experience to localize failures across a layered data-center architecture spanning switching/forwarding, services/enforcement, and control-plane domains - Linux operations experience at the command line, including production troubleshooting, with working exposure to containers or Kubernetes ## Nice to Have - Direct experience using network troubleshooting tooling as a primary diagnostic method, including packet capture and flow-telemetry analysis (NetFlow/IPFIX) - Working knowledge of enterprise virtualization, sufficient to troubleshoot a VM-based appliance deployment; vSphere is the current deployment target - Operational Kubernetes and Helm proficiency, including diagnosing failures beyond the workload level: TLS certificates and service-account authentication, API-server connectivity, service exposure, persistent storage for stateful workloads, and custom resources and operators - Working knowledge of VXLAN EVPN fabrics, including the Smart Switch's placement within them and the ability to isolate faults across the fabric - Network segmentation and firewall policy design (zone-based or microsegmentation), and the tradeoffs vs. traditional NGFWs - Familiarity with the NetOps/NetSecOps operating split in data-center security - Experience driving diagnosis and remediation through a customer's own team, working incidents in environments with no direct access where the customer accomplishes the steps - Experience with NX-OS automation and APIs (NX-API, NETCONF/RESTCONF, gNMI, or Ansible); familiarity with Cisco Nexus Dashboard a plus - Demonstrated ability to communicate incident status, root cause, and remediation clearly to both technical and executive audiences, verbally and in writing - CCNP Data Center, CCNP Enterprise, DevNet Professional, CCIE Data Center, CCIE Enterprise, or DevNet Expert (a plus) ## Benefits At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations in the AI era – and beyond. We’ve been innovating fearlessly for 40 years to create solutions that power how humans and technology work together across the physical and digital worlds. These solutions provide customers with unparalleled security, visibility, and insights across the entire digital footprint. Fueled by the depth and breadth of our technology, we experiment and create meaningful solutions. Add to that our worldwide network of doers and experts, and you’ll see that the opportunities to grow and build are limitless. We work as a team, collaborating with empathy to make really big things happen on a global scale. Because our solutions are everywhere, our impact is everywhere. We are Cisco, and our power starts with you.
Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗