AI Networking
Core
Design, bring up, and scale distributed Ethernet and InfiniBand fabrics connecting hundreds of thousands of GPUs across multi-megawatt data halls.
Role type
Senior IC AI Networking Engineer (High-Performance Computing)
Builds
Distributed GPU fabrics, ROCE transport networks, and AI training/inference clusters
Domain
High-performance computing, AI infrastructure, data center networking
Deliverable
production ML models | infrastructure
Required skills
Distributed system design, ROCE transport design, congestion control tuning, network modeling, telemetry engineering, automated troubleshooting, routing algorithm development, performance benchmarking, root-cause analysis, C/C++/Python/Java/JavaScript/C# programming
Preferred skills
Experience with NVIDIA/Broadcom network designs, silicon/network co-design, pretraining compute roadmap development
Technologies
Ethernet, InfiniBand, ROCE, ECN, WRED, DCTCP
Responsibilities
Design and scale distributed Ethernet and InfiniBand fabrics for GPU clusters; tune ROCE transport and congestion control mechanisms; develop novel routing techniques for large-scale network reliability; perform AI cluster bring-up and performance benchmarking; gather data to develop pretraining compute roadmaps
Seniority
Senior, hands-on IC
