大模型推理引擎研发工程师(深圳/北京/上海/杭州)
Core
Develop and optimize large language model (LLM) inference engines and PD-separated inference scheduling systems to maximize performance and cost efficiency.
Role type
Senior IC LLM inference engine engineer
Builds
High-performance LLM inference systems supporting mainstream GPUs and heterogeneous AI chips
Domain
Artificial Intelligence / Large Language Models / GPU Computing
Deliverable
production ML models
Required skills
C/C++, Python, CUDA, OpenCL, Ascend C, Cutlass, vLLM, SGLang, TensorRT-LLM, FasterTransformer, deep learning operator implementation, model parallelism, pipeline parallelism, NVLINK, GPU communication, GPU/AI chip architecture, system performance analysis
Preferred skills
Top-tier conference publications in machine learning or computer architecture, open-source contributions to vLLM or SGLang, experience with large-scale model distributed deployment, inference service framework deployment
Technologies
CUDA, OpenCL, Ascend C, Cutlass, vLLM, SGLang, TensorRT-LLM, FasterTransformer, NVLINK