Network Reliability Engineer III (Hybrid, IND)
Core
Implement, maintain, and optimize AI-powered tools and platforms for network monitoring, alerting, and incident response to enhance reliability in hyper-scale hybrid cloud infrastructure.
Role type
Senior Network Reliability Engineer (AI/ML integration)
Builds
AI-driven anomaly detection systems, automated remediation workflows, and intelligent alerting platforms
Domain
Cybersecurity, Network Engineering, AI/ML Operations
Deliverable
production ML models | infrastructure
Required skills
Large scale IP networking (BGP, OSPF, ISIS, VRF, VxLAN, eVPN, QoS, GRE, IPSec, DNS, MACsec, MPLS, CLOS), Linux troubleshooting, observability tools (Prometheus, Grafana, ELK), Infrastructure-as-code (Python, Terraform, Ansible, YAML, Netconf/YANG, Chef, Puppet), AI/ML platform integration, network telemetry analysis
Preferred skills
Data center/telecom/SaaS/cloud operations experience, device lifecycle management, AI-powered monitoring platforms (Moogsoft, BigPanda, Splunk, DataDog), cloud ML services (AWS/Azure/GCP), ticketing systems (ServiceNow, Jira)
Technologies
BGP, OSPF, ISIS, VRF, VxLAN, eVPN, QoS, GRE, IPSec, DNS, MACsec, MPLS, CLOS, Prometheus, Grafana, ELK, Python, Terraform, Ansible, YAML, Netconf, YANG, Chef, Puppet, Moogsoft, BigPanda, Splunk, DataDog, AWS, Azure, GCP, ServiceNow, Jira
Responsibilities
Deploy and configure AI-powered network monitoring and analytics platforms; Implement and maintain AI-driven anomaly detection systems; Configure and optimize automated remediation workflows; Ensure network telemetry integration with ML platforms; Fine-tune intelligent alerting systems; Utilize smart correlation tools for incident detection; Support and maintain AI-enhanced infrastructure; Create documentation for automated procedures; Collaborate with Network Engineering and TechOps teams; Participate in 24x7 on-call rotation
Seniority
Senior, hands-on IC