Senior Network Site Reliability Engineer
Core
Staff Network SRE ensuring high availability and reliability of enterprise network infrastructure through hands-on debugging, automation, and observability.
Role type
Staff Network Site Reliability Engineer
Builds
Reliable and efficient network infrastructure for enterprise users
Domain
Enterprise Networking & Data Center Operations
Deliverable
production ML models | infrastructure
Required skills
Network fundamentals (TCP/UDP, IPv4/IPv6, BGP, OSPF, ISIS, VPN, L2 switching), Network automation (Salt, Ansible, Python), Linux system administration, Monitoring tools (Prometheus, Grafana, Alert Manager, Nautobot/Netbox, BigPanda), Service management (ServiceNow, Jira, ITIL), Root Cause Analysis (RCA)
Preferred skills
Streaming Telemetry (SNMP, Syslog), Advanced network technologies (VXLAN/EVPN, MPLS, RSVP, Segment Routing, SDWAN, SASE), Programming in Python or Go, Experience with specific hardware (Mellanox/Cumulus, Cisco/Arista, Palo Alto, Versa, Netscalers, F5)
Responsibilities
Own operational aspects of network infrastructure ensuring high availability, Partner with architecture teams to ensure supportable implementations, Implement automation to reduce toil and maintain SLOs, Monitor network performance and mitigate risks, Conduct blameless postmortems and RCAs, Develop knowledge base articles for automation
Seniority
Staff, hands-on IC with strategic impact