Senior Network & Site Reliability Engineer
Core
Design and operate the global network and reliability layer for one of the world's fastest private supercomputers (NVIDIA DGX SuperPOD) powering distributed compute and ML workloads.
Role type
Senior Network & Site Reliability Engineer
Builds
Scalable, secure network architecture and reliability infrastructure for high-performance distributed systems
Domain
Data Center Networking / High-Performance Computing / Machine Learning Infrastructure
Deliverable
infrastructure
Required skills
Network architecture design, Network device configuration management, Network automation and IaC, WAN engineering, Kubernetes networking, Linux system administration, Monitoring and observability, Python/Bash scripting
Preferred skills
NVIDIA networking technologies (Cumulus Linux, InfiniBand, Spectrum-X), Data-intensive platform experience, High-compliance environment experience
Technologies
Ansible, Terraform, Nornir, NetBox, Infoblox, Prometheus, Grafana, Datadog, ELK, OpenTelemetry, Kubernetes, BGP, MPLS, IPsec VPN, InfiniBand, Spectrum-X, BlueField
Responsibilities
Architect and operate scalable, secure network architecture for large-scale ML workloads; Own network device configuration management end to end; Improve system and network reliability through automation and proactive capacity planning; Implement and manage complex network protocols (BGP, VPNs, WAN circuits); Build and maintain monitoring, alerting, and incident response systems; Ensure security and compliance across network infrastructure; Partner with engineering and data science teams.
Seniority
Senior, hands-on IC