北京-AI infra 推理工程师(基座研发方向)(J101235)
Core
Designing and evolving the inference infrastructure for Baidu's Wenxin large language models, focusing on GPU-based performance optimization and building a self-reliant high-performance inference base.
Role type
Senior IC AI infrastructure inference engineer (LLM base research direction)
Builds
Autonomous high-performance inference engines, high-performance computing libraries, and communication libraries for large models.
Domain
AI Infrastructure / Large Language Models / High-Performance Computing
Deliverable
production ML models
Required skills
C++, CUDA, Python, parallel computing, distributed systems, attention mechanism optimization, matrix computation optimization, batch scheduling strategies, cache pooling, tensor core optimization
Preferred skills
vLLM, SGLang, TensorRT-LLM, speculative decoding, sparse attention, low-bit quantization, agent scenario cache optimization, top-tier conference publications (OSDI/SOSP/MLSys/NeurIPS/ICML)
Technologies
CUDA, CUTLASS, CuTe, Triton, TileLang, vLLM, SGLang, TensorRT-LLM
Responsibilities
Develop and optimize the full-stack inference capabilities including functional development, post-training adaptation, and deep performance tuning; Design and iterate inference engines and communication libraries; Collaborate on model structure and inference optimization; Track and validate frontier AI infrastructure technologies.
Seniority
Senior, hands-on IC