CareerPlanSign in

Senior Network & Site Reliability Engineer

San Francisco HQ💼 Full-time🗓 2026-06-12 → 2026-09-26

Core

Design and operate the global network and reliability layer for one of the world's fastest private supercomputers (NVIDIA DGX SuperPOD) powering distributed compute and ML workloads.

Role type

Senior Network & Site Reliability Engineer

Builds

Scalable, secure network architecture and reliability infrastructure for high-performance distributed systems

Domain

Data Center Networking / High-Performance Computing / Machine Learning Infrastructure

Deliverable

infrastructure

Required skills

Network architecture design, Network device configuration management, Network automation and IaC, WAN engineering, Kubernetes networking, Linux system administration, Monitoring and observability, Python/Bash scripting

Preferred skills

NVIDIA networking technologies (Cumulus Linux, InfiniBand, Spectrum-X), Data-intensive platform experience, High-compliance environment experience

Technologies

Ansible, Terraform, Nornir, NetBox, Infoblox, Prometheus, Grafana, Datadog, ELK, OpenTelemetry, Kubernetes, BGP, MPLS, IPsec VPN, InfiniBand, Spectrum-X, BlueField

Responsibilities

Architect and operate scalable, secure network architecture for large-scale ML workloads; Own network device configuration management end to end; Improve system and network reliability through automation and proactive capacity planning; Implement and manage complex network protocols (BGP, VPNs, WAN circuits); Build and maintain monitoring, alerting, and incident response systems; Ensure security and compliance across network infrastructure; Partner with engineering and data science teams.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.