CareerPlanSign in

北京-AI infra 推理工程师(基座研发方向)(J101235)

北京市💼 Full-time🗓 2026-07-21 → 2026-09-28

Core

Designing and evolving the inference infrastructure for Baidu's Wenxin large language models, focusing on GPU-based performance optimization and building a self-reliant high-performance inference base.

Role type

Senior IC AI infrastructure inference engineer (LLM base research direction)

Builds

Autonomous high-performance inference engines, high-performance computing libraries, and communication libraries for large models.

Domain

AI Infrastructure / Large Language Models / High-Performance Computing

Deliverable

production ML models

Required skills

C++, CUDA, Python, parallel computing, distributed systems, attention mechanism optimization, matrix computation optimization, batch scheduling strategies, cache pooling, tensor core optimization

Preferred skills

vLLM, SGLang, TensorRT-LLM, speculative decoding, sparse attention, low-bit quantization, agent scenario cache optimization, top-tier conference publications (OSDI/SOSP/MLSys/NeurIPS/ICML)

Technologies

CUDA, CUTLASS, CuTe, Triton, TileLang, vLLM, SGLang, TensorRT-LLM

Responsibilities

Develop and optimize the full-stack inference capabilities including functional development, post-training adaptation, and deep performance tuning; Design and iterate inference engines and communication libraries; Collaborate on model structure and inference optimization; Track and validate frontier AI infrastructure technologies.

Seniority

Senior, hands-on IC

Sourced via baidu · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.