元宝-LLM大模型推理工程师
Core
Deploy, operate, and optimize LLM inference services for business scenarios, focusing on acceleration and integration of advanced inference techniques.
Role type
LLM inference optimization engineer
Builds
Optimized LLM inference services for search and business applications
Domain
Artificial Intelligence / Large Language Models
Deliverable
production ML models
Required skills
C++, Python, Go, GPU programming (CUDA, OpenCL), inference frameworks (TensorRT, Triton, SGLang, vLLM), model quantization, pruning, dynamic batching, operator fusion, distributed inference
Preferred skills
Large-scale model distributed deployment experience
Technologies
CUDA, OpenCL, cuBLAS, cuDNN, TensorRT, Triton, SGLang, vLLM
Responsibilities
Implement and deploy inference acceleration methods (pruning, quantization, dynamic batch); Research and integrate frontier techniques (sparse, heterogeneous, distributed inference) into search business; Optimize LLM deployment and operations for service scenarios.