腾讯游戏-大模型推理性能优化工程师/专家
Core
Design and optimize high-performance inference engines for LLM/VLM/DiT models to enable efficient deployment and maximize cost-performance.
Role type
Senior IC machine-learning inference optimization engineer
Builds
High-performance inference engines for large-scale distributed systems
Domain
AI/ML inference optimization on GPU and heterogeneous AI chips
Deliverable
production ML models
Required skills
C/C++, Python, large model inference frameworks (vllm, sglang, tensorrt-llm), parallel strategies (data/pipe-parallelism), GPU/AI chip architecture, deep learning operator implementation
Preferred skills
NVLINK/GPU RDMA communication, system performance analysis and tuning, open-source model architecture analysis
Technologies
vllm, sglang, tensorrt-llm, NVLINK, GPU RDMA
Responsibilities
Collaborate with algorithm teams to build industry-leading inference engines; Optimize inference performance via PD separation, low-bit computation, and parallelism; Support mainstream GPUs and heterogeneous AI chips for cost-effective deployment.