北京-AI通信研发工程师(J104693)
Core
Optimize cross-node communication efficiency in large-scale clusters and tune intra-node communication for large model inference performance.
Role type
Senior IC distributed systems engineer (RDMA & high-performance networking)
Builds
High-performance distributed training and inference clusters for large language models
Domain
AI infrastructure, high-performance networking, distributed systems
Deliverable
production ML models
Required skills
RDMA protocol stack (RoCEv2, InfiniBand), C/C++, Python, Linux network programming, DPDK, congestion control algorithms (DCQCN, PFC, ECN), distributed parallelism strategies (TP, PP, EP), vLLM/SGLang frameworks
Preferred skills
P4/eBPF programmable NICs, ns-3 network simulation, Transformer architecture, open-source contributions
Technologies
RDMA, RoCEv2, InfiniBand, DPDK, vLLM, SGLang, C++, Python, Linux
Responsibilities
Analyze and optimize RDMA protocol stack performance in real-world scenarios; design and execute benchmark test suites for RDMA NICs; research and implement adaptive congestion control algorithms on programmable NICs; tune communication logic for tensor/pipeline/expert parallelism in large model inference
Seniority
Mid-Senior, hands-on IC
