CareerPlanSign in

大模型推理引擎研发工程师(深圳/北京/上海/杭州)

Beijing, China💼 Full-time🗓 2026-09-28

Core

Develop and optimize large language model (LLM) inference engines and PD-separated inference scheduling systems to maximize performance and cost efficiency.

Role type

Senior IC LLM inference engine engineer

Builds

High-performance LLM inference systems supporting mainstream GPUs and heterogeneous AI chips

Domain

Artificial Intelligence / Large Language Models / GPU Computing

Deliverable

production ML models

Required skills

C/C++, Python, CUDA, OpenCL, Ascend C, Cutlass, vLLM, SGLang, TensorRT-LLM, FasterTransformer, deep learning operator implementation, model parallelism, pipeline parallelism, NVLINK, GPU communication, GPU/AI chip architecture, system performance analysis

Preferred skills

Top-tier conference publications in machine learning or computer architecture, open-source contributions to vLLM or SGLang, experience with large-scale model distributed deployment, inference service framework deployment

Technologies

CUDA, OpenCL, Ascend C, Cutlass, vLLM, SGLang, TensorRT-LLM, FasterTransformer, NVLINK

Sourced via tencent · Listed on CareerPlan, which tracks 850,000+ jobs from 20+ sources.