CareerPlanSign in

元宝-LLM大模型推理工程师

Shenzhen, China💼 Full-time🗓 2026-09-28

Core

Develop and optimize large language model (LLM) inference frameworks and high-performance inference engines.

Role type

Senior IC LLM inference optimization engineer

Builds

High-performance LLM inference engines

Domain

AI/ML infrastructure, GPU computing

Deliverable

production ML models

Required skills

GPU high-performance computing optimization, CUDA programming, computer architecture, parallel computing, memory optimization, low-bit computation, deep learning framework internals (PyTorch/TensorFlow), LLM model acceleration techniques (subgraph matching, compilation, quantization)

Preferred skills

Experience with TensorRT-LLM, vLLM, AI engineering optimization

Technologies

CUDA, GPU, TensorRT-LLM, vLLM, PyTorch, TensorFlow

Responsibilities

Develop and optimize LLM inference frameworks; Optimize high-performance LLM inference engines using GPU and CUDA; Research and introduce forward-looking technical architectures for LLM training and inference.

Seniority

Senior, hands-on IC

Sourced via tencent · Listed on CareerPlan, which tracks 848,000+ jobs from 20+ sources.