Staff Reliability Engineer
Core
Design and own the resilience, health, and availability of the global corporate network, ensuring an 'Always Secure. Always On.' environment for employees worldwide.
Role type
Staff Reliability Engineer (Network Operations)
Builds
Scalable self-service operational tooling, automated network operations, and resilient global network infrastructure.
Domain
Enterprise Networking, Cloud Infrastructure, Distributed Systems
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
AWS Networking, Palo Alto Networks solutions, Distributed Systems fundamentals, Networking fundamentals (WiFi, DNS, DHCP, VLANs, VPN, ACLs, Routing, Firewall Policies), Infrastructure as Code (Terraform/Ansible), Observability (Prometheus/Grafana), Python/Go programming, Service Reliability Management (SLOs/SLIs)
Preferred skills
Juniper/JUNOS switching/routing, Palo Alto Networks NGFWs, enterprise office build and construction processes
Technologies
AWS, Palo Alto Networks, Terraform, Ansible, Prometheus, Grafana, Python, Go, Juniper, JUNOS
Responsibilities
Respond to alerts and monitor system health to ensure network availability; Drive strategic reduction of systemic toil and technical debt through automation; Collaborate with cross-functional stakeholders to resolve complex network operations issues; Mentor team members and define success metrics for the team; Lead the resolution of complex network operations issues from a systems perspective.
Seniority
Staff, strategic ownership & mentorship
