HPC Network Engineer
Core
Design, deploy, and operate the network fabric for a multi-tenant AI cluster, covering high-performance compute/storage fabrics, tenant isolation, and office networks.
Role type
Senior HPC Network Engineer (AI Cluster)
Builds
Lossless RDMA fabrics (RoCEv2, InfiniBand), leaf-spine data center networks, and multi-tenant isolation for GPU workloads.
Domain
High-Performance Computing / AI Infrastructure
Deliverable
production ML models | infrastructure
Required skills
RDMA fabric design (RoCEv2, InfiniBand), dynamic routing (BGP, EVPN/VXLAN), leaf-spine/Clos architecture, network automation (Python, Ansible), Linux networking stack, network telemetry/observability, corporate network management.
Preferred skills
GPU cluster networking (NVIDIA Spectrum-X, ConnectX/BlueField), container networking (CNI, BGP), collective-communication libraries, storage networking (NVMe-oF), optical layer knowledge (200/400/800G).
Technologies
RoCEv2, InfiniBand, BGP, EVPN, VXLAN, Ansible, Python, NetBox, Prometheus, Grafana, Datadog, Cisco Catalyst, Meraki, FortiGate, Palo Alto.
Responsibilities
Design and operate lossless RDMA-capable fabrics for GPU compute and storage; build and manage leaf-spine data centre fabrics with routed underlay and overlay; implement per-tenant network isolation; automate network provisioning and configuration; build telemetry and observability for the fabric; troubleshoot performance issues end to end; operate the out-of-band management network; support tenant onboarding; own and maintain the office network; upskill colleagues on networking.
Seniority
Senior, hands-on IC
