CareerPlanSign in

元宝-LLM大模型推理工程师

Beijing, China💼 Full-time🗓 2026-09-28

Core

Deploy, operate, and optimize LLM inference services for business scenarios, focusing on acceleration and integration of advanced inference techniques.

Role type

LLM inference optimization engineer

Builds

Optimized LLM inference services for search and business applications

Domain

Artificial Intelligence / Large Language Models

Deliverable

production ML models

Required skills

C++, Python, Go, GPU programming (CUDA, OpenCL), inference frameworks (TensorRT, Triton, SGLang, vLLM), model quantization, pruning, dynamic batching, operator fusion, distributed inference

Preferred skills

Large-scale model distributed deployment experience

Technologies

CUDA, OpenCL, cuBLAS, cuDNN, TensorRT, Triton, SGLang, vLLM

Responsibilities

Implement and deploy inference acceleration methods (pruning, quantization, dynamic batch); Research and integrate frontier techniques (sparse, heterogeneous, distributed inference) into search business; Optimize LLM deployment and operations for service scenarios.

Sourced via tencent · Listed on CareerPlan, which tracks 853,000+ jobs from 20+ sources.