Senior Network Reliability Engineer - DGX Cloud
Core
Support and maintain cloud and datacenter network infrastructures serving NVIDIA's software stack from Graphics Drivers to Autonomous Vehicles and AI.
Role type
Senior Network Reliability Engineer
Builds
Cloud and datacenter network infrastructure for NVIDIA's software stack
Domain
Cloud computing, Datacenter networking, AI infrastructure
Deliverable
infrastructure
Required skills
TCP/IP, BGP, OSPF, MPLS, IS-IS, VxLAN, EVPN, QoS, GRE, IPsec, DNS, MACsec, network troubleshooting, incident management, alert response, CSP environment experience (AWS, Azure, GCP, OCI), network device management (Arista, Fortinet, Juniper), automation scripting (Python/Shell)
Preferred skills
Mellanox/Cumulus OS, Infiniband technology, Unix/Linux system administration, monitoring tools (Netbox/Nautobot, Prometheus, Grafana, Panoptes)
Technologies
PNI, Transit, Exchange, Passive DWDM, Wave circuits, Mellanox, Infiniband, Netbox, Nautobot, Prometheus, Grafana, Panoptes
Responsibilities
Remediate critical alerts within SLAs, triage production impacting network incidents, engage with external vendors for hardware/software issues, participate in network device upgrades and capacity augmentations, manage large scale IP network technologies, monitor network health of on-premises and cloud infrastructures
Seniority
Senior, hands-on IC