Staff Network Reliability Engineer (f/m/d)
Core
Operate large-scale production data center networks (physical fabric, SDN overlay, virtual network functions) with a reliability mindset, defining health metrics and leading incident response.
Role type
Staff Network Reliability Engineer
Builds
Reliable data center network infrastructure across dozens of locations
Domain
Data Center Networking / Network Operations
Deliverable
production ML models | infrastructure
Required skills
Large-scale data center network operations, Layer 2/3 routing (BGP, OSPF), VXLAN/BGP-EVPN, Multi-vendor operations (Juniper, Cisco), Linux fundamentals, Python, Ansible, CI/CD, SLO definition, Incident command, Observability design, AI application for anomaly detection and automated diagnosis
Preferred skills
SONiC, Netbox, German language
Technologies
Juniper, Cisco, SONiC, Netbox, Python, Ansible, GitLab, bash, tcpdump, iptables, Git
Responsibilities
Own reliability of physical fabric, SDN overlay, and virtual network functions; Lead incident response and postmortems; Remove toil through automation pipelines; Apply AI to operations for anomaly detection and automated remediation; Strengthen observability and shape lifecycle practices; Collaborate with provisioning teams; Share on-call rotation
Seniority
Staff, hands-on IC with strategic influence