Senior AI Network Systems Engineer
Core
Define and develop networking requirements for large-scale AI training and inference clusters, delivering scalable networking solutions from concept through datacenter deployment.
Role type
Senior IC AI Network Systems Engineer
Builds
Scalable networking solutions for large-scale AI training and inference clusters
Domain
AI infrastructure, high-performance computing, cloud data centers
Deliverable
production ML models | infrastructure
Required skills
RDMA technologies, AI fabrics, distributed training environments, RoCE, congestion control, ECN, PFC, DCQCN, AI/ML workload communication patterns, collective operations, SONiC, Linux networking, network telemetry, network operating systems, network switches, SmartNICs, DPUs, NIC offloads, large-scale cloud infrastructure, network stress tools, validation frameworks, performance benchmarks, observability solutions, packet analysis tools, telemetry infrastructure, network automation frameworks, high-speed networking environments (200G/400G/800G Ethernet)
Preferred skills
Experience developing network stress tools, validation frameworks, performance benchmarks, or observability solutions
Technologies
RDMA, RoCE, ECN, PFC, DCQCN, SONiC, Linux, 200G/400G/800G Ethernet
Responsibilities
Define network concepts of operation, serviceability requirements, telemetry requirements, and operational models for AI infrastructure; Analyze transport-layer behavior and performance characteristics across large-scale distributed AI workloads; Design, validate, and optimize RDMA-based networking solutions for AI clusters; Develop and execute networking validation strategies covering functionality, performance, scale, interoperability, resiliency, and reliability; Perform packet-level analysis and protocol debugging using telemetry, packet captures, performance counters, and diagnostic tools; Build and improve network observability, diagnostics, telemetry, and monitoring solutions
Seniority
Senior, hands-on IC