CareerPlanGet AI match score →

Senior AI Network Systems Engineer

United States, Washington, Redmond💼 Full-time🗓 2026-07-16 → 2026-07-31

Core

Define and develop networking requirements for large-scale AI training and inference clusters, delivering scalable networking solutions from concept through datacenter deployment.

Role type

Senior IC AI Network Systems Engineer

Builds

Scalable networking solutions for large-scale AI training and inference clusters

Domain

AI infrastructure, high-performance computing, cloud data centers

Deliverable

production ML models | infrastructure

Required skills

RDMA technologies, AI fabrics, distributed training environments, RoCE, congestion control, ECN, PFC, DCQCN, AI/ML workload communication patterns, collective operations, SONiC, Linux networking, network telemetry, network operating systems, network switches, SmartNICs, DPUs, NIC offloads, large-scale cloud infrastructure, network stress tools, validation frameworks, performance benchmarks, observability solutions, packet analysis tools, telemetry infrastructure, network automation frameworks, high-speed networking environments (200G/400G/800G Ethernet)

Preferred skills

Experience developing network stress tools, validation frameworks, performance benchmarks, or observability solutions

Technologies

RDMA, RoCE, ECN, PFC, DCQCN, SONiC, Linux, 200G/400G/800G Ethernet

Responsibilities

Define network concepts of operation, serviceability requirements, telemetry requirements, and operational models for AI infrastructure; Analyze transport-layer behavior and performance characteristics across large-scale distributed AI workloads; Design, validate, and optimize RDMA-based networking solutions for AI clusters; Develop and execute networking validation strategies covering functionality, performance, scale, interoperability, resiliency, and reliability; Perform packet-level analysis and protocol debugging using telemetry, packet captures, performance counters, and diagnostic tools; Build and improve network observability, diagnostics, telemetry, and monitoring solutions

Seniority

Senior, hands-on IC

Rewrite
## About the role * Define and develop networking requirements for large-scale AI training and inference clusters. * Collaborate with silicon, system software, firmware, hardware, and Azure infrastructure teams to deliver scalable networking solutions from concept through datacenter deployment. * Participate in architecture reviews and influence next-generation AI networking roadmaps. * Define network concepts of operation, serviceability requirements, telemetry requirements, and operational models for AI infrastructure. * Analyze transport-layer behavior and performance characteristics across large-scale distributed AI workloads. * Evaluate network protocol implementations and debug issues impacting latency, throughput, scalability, and reliability. * Design, validate, and optimize RDMA-based networking solutions for AI clusters. * Analyze RDMA performance, congestion behavior, packet loss, retransmissions, and collective communication efficiency. * Work closely with networking vendors and software teams to optimize AI fabric performance and workload scalability. * Develop validation methodologies for AI traffic patterns and collective communication workloads. ## Performance Characterization & Validation * Develop and execute networking validation strategies covering functionality, performance, scale, interoperability, resiliency, and reliability. * Evaluate latency, bandwidth utilization, congestion events, flow distribution, and workload communication patterns. * Perform packet-level analysis and protocol debugging using telemetry, packet captures, performance counters, and diagnostic tools. * Investigate network switch, NIC, RDMA, routing, congestion control, and protocol-related issues. * Build and improve network observability, diagnostics, telemetry, and monitoring solutions. * Develop tools and automation for network validation, performance analysis, and failure detection. * Improve engineering productivity through automated testing, qualification, and network health assessment frameworks. ## Requirements * Master's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 3+ years technical engineering experience * OR Bachelor's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 5+ years technical engineering experience * OR equivalent experience. * 5+ years of experience developing or validating networking for accelerator based systems. * 5+ years of experience designing, integrating, validating, or troubleshooting Ethernet-based networking infrastructure, including Network switches. * 5+ years of experience supporting AI, HPC, cloud, or large-scale data center infrastructure deployments. ## Nice to have * These requirements include but are not limited to the following specialized security screenings: * Experience with RDMA technologies, AI fabrics, and distributed training environments. * Understanding of RoCE, congestion control, ECN, PFC, DCQCN, and related AI networking technologies. * Experience with AI/ML workload communication patterns and collective operations. * Experience with SONiC, Linux networking, networking telemetry, and network operating systems. * Experience with network switches, SmartNICs, DPUs, NIC offloads, and large-scale cloud infrastructure. * Familiarity with AI networking technologies including Ultra Ethernet and hyperscale AI cluster architectures. * Experience developing network stress tools, validation frameworks, performance benchmarks, or observability solutions. * Knowledge of packet analysis tools, telemetry infrastructure, and network automation frameworks. * Exposure to high-speed networking environments (200G/400G/800G Ethernet).
Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply at Microsoft ↗