Operations Engineer
Core
Ensure the reliability, availability, and performance of enterprise network services through automation, proactive monitoring, and continuous improvement.
Role type
Senior Network Site Reliability Engineer (SRE)
Builds
Automated network solutions, self-healing capabilities, and resilient service architectures
Domain
Network Infrastructure & Site Reliability Engineering
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Network automation, Python, SD-WAN, Routing & Switching (BGP, OSPF, MPLS), Network Observability, Incident & Problem Management, Infrastructure as Code, Log Analytics, Capacity Planning
Preferred skills
Cisco DevNet, SASE/Prisma Access, Genesys & Oracle SBC, Auto-remediation strategies
Technologies
Cisco, Juniper, Arista, Genesys, Oracle SBC, Python, IaC tools
Responsibilities
Develop and maintain network automation solutions to reduce manual effort and operational risk. Drive proactive monitoring, alerting, and incident prevention initiatives. Lead and support major incident resolution and problem management activities. Implement auto-remediation and self-healing capabilities for network services. Perform capacity planning, performance analysis, and trend forecasting. Collaborate with Network, Security, Cloud, and Application teams to improve end-to-end service stability. Create and maintain SOPs, runbooks, and operational documentation. Analyze logs, telemetry, and monitoring data to identify recurring issues and technical debt. Mentor operations teams and promote SRE best practices across the network organization.
Seniority
Senior, hands-on IC