CareerPlanSign in

腾讯游戏-大模型推理性能优化工程师/专家

Shanghai, China💼 Full-time🗓 2026-09-28

Core

Design and optimize high-performance inference engines for LLM/VLM/DiT models to enable efficient deployment and maximize cost-performance.

Role type

Senior IC machine-learning inference optimization engineer

Builds

High-performance inference engines for large-scale distributed systems

Domain

AI/ML inference optimization on GPU and heterogeneous AI chips

Deliverable

production ML models

Required skills

C/C++, Python, large model inference frameworks (vllm, sglang, tensorrt-llm), parallel strategies (data/pipe-parallelism), GPU/AI chip architecture, deep learning operator implementation

Preferred skills

NVLINK/GPU RDMA communication, system performance analysis and tuning, open-source model architecture analysis

Technologies

vllm, sglang, tensorrt-llm, NVLINK, GPU RDMA

Responsibilities

Collaborate with algorithm teams to build industry-leading inference engines; Optimize inference performance via PD separation, low-bit computation, and parallelism; Support mainstream GPUs and heterogeneous AI chips for cost-effective deployment.

Sourced via tencent · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.