大模型推理架构研发工程师(J95970)
Core
Optimize inference performance for Baidu's ERNIE Bot large language models and design/develop the PaddlePaddle deep learning inference framework.
Role type
Senior IC machine-learning infrastructure engineer (LLM inference)
Builds
High-performance inference frameworks, heterogeneous computing platforms, and communication libraries for AI models.
Domain
Artificial Intelligence / Large Language Models / High-Performance Computing
Deliverable
production ML models
Required skills
C++, Python, CUDA, computer architecture, assembly-level development, FlashAttention, PagedAttention, MoE, Chunked Prefill, quantization algorithms (AWQ, GPTQ, SmoothQuant), communication operators (Allreduce), compute-communication overlap, separated deployment (PD separation)
Preferred skills
PaddlePaddle, PyTorch, TensorFlow, vLLM, TGI, SGLang, TensorRT-LLM
Responsibilities
Design and develop heterogeneous high-performance computing platforms; optimize deep learning frameworks on CPU/GPU; track and implement cutting-edge LLM technologies; support business units like autonomous driving and search with inference optimization.
Seniority
Senior, hands-on IC