CareerPlanSign in

腾讯游戏-大模型推理性能优化工程师/专家

Hangzhou, China💼 Full-time🗓 2026-09-28

Core

Design and optimize high-performance inference engines for Large Language Models (LLM), Vision-Language Models (VLM), and Diffusion Transformers (DiT) to enable efficient deployment.

Role type

Senior IC machine-learning inference optimization engineer

Builds

High-performance distributed inference systems for LLM/VLM/DiT models

Domain

AI/ML inference optimization on GPU and heterogeneous AI chips

Deliverable

production ML models

Required skills

C/C++, Python, distributed inference strategies (data/pipe-parallelism), GPU architecture, deep learning operator optimization, model architecture analysis

Preferred skills

NVLINK/GPU RDMA communication, low-bit quantization, PD separation, heterogeneous AI chip tuning

Technologies

vllm, sglang, tensorrt-llm

Responsibilities

Collaborate with algorithm teams to build industry-leading inference engines; Optimize inference performance via parallelism, low-bit computation, and operator tuning; Support mainstream GPUs and AI chips to maximize performance-cost efficiency.

Sourced via tencent · Listed on CareerPlan, which tracks 902,000+ jobs from 20+ sources.