Senior Network Reliability Engineer AI - Data Center / BGP (m/w/d)
Core
Own reliability of physical fabric, SDN overlay, and virtual network functions across data centers, setting SLOs and driving down unplanned downtime.
Role type
Senior IC network reliability engineer (AI/automation)
Builds
Reliable data center networks with AI-driven operations and automation pipelines
Domain
Cloud services and hosting / Data center networking
Deliverable
production ML models | infrastructure
Required skills
Large-scale production data center network operations, Layer 2/3 routing (BGP, OSPF), VXLAN/BGP-EVPN, multi-vendor operations (Juniper, Cisco), Linux fundamentals (bash, tcpdump, iptables, git), Python automation (Ansible, CI/CD), reliability engineering (SLOs, incident command, postmortems), AI application in network operations
Preferred skills
SONiC, Netbox, German language
Technologies
BGP, OSPF, VXLAN, EVPN, Juniper, Cisco, SONiC, Netbox, Python, Ansible, GitLab, bash, tcpdump, iptables, git
Responsibilities
Own reliability of physical fabric, SDN overlay, and virtual network functions; lead incident response on major network incidents and run postmortems; remove toil through automation pipelines and tooling; apply AI to operations for anomaly detection, diagnosis, config generation, and remediation; strengthen observability and shape lifecycle, upgrade, and capacity practices
Seniority
Senior, hands-on IC