Network Engineer
Core
Design and architect front-end network fabrics for large-scale AI/ML and HPC clusters, optimizing for high resource utilization, low latency, and high-throughput communication.
Role type
Senior Network Architect (AI Cluster Infrastructure)
Builds
Resilient, reliable, high-throughput interconnect fabrics for AI training and inference clusters
Domain
AI/ML infrastructure, High-Performance Computing (HPC), Datacenter Networking
Deliverable
production ML models | infrastructure
Required skills
Large-scale datacenter network design, Distributed systems debugging, Python/Go programming, Streaming telemetry (gNMI, OpenConfig), RoCEv2, VXLAN, EVPN, BGP, SRE practices
Preferred skills
Hyperscaler experience, AI/ML cluster networking, Open-source networking contributions
Technologies
Juniper, Arista, Cisco, SONiC, Prometheus, InfluxDB, Grafana, Ansible, Jinja2, CI/CD pipelines
Responsibilities
Design front-end network fabrics for AI/ML clusters, Build proof-of-concept implementations of new network designs, Automate network infrastructure deployment and validation, Operate SRE-grade telemetry and observability, Lead network debugging in large distributed systems, Represent the company in industry forums and standards bodies
Seniority
Senior, hands-on IC with architectural leadership