CareerPlanSign in

微信搜索-AI Infra 工程师-大模型推理方向 (深圳)(广州)

Beijing, China💼 Full-time🗓 2026-09-28

Core

Design and optimize inference engines for Large Language Models (LLM) and Vision-Language Models (VLM) to support AI Search and intelligent Agent applications.

Role type

Senior IC LLM Infrastructure Engineer (Inference)

Builds

Inference infrastructure for large-scale AI Search and Agent systems

Domain

AI / Large Language Models / Search Technology

Deliverable

production ML models

Required skills

LLM/VLM model architecture, inference engine optimization (vllm/sglang/TRT-llm), operator fusion, quantization strategies, dynamic batching, distributed KV cache optimization, large-scale system design

Preferred skills

Experience with AI hardware configuration, model structure optimization for real-world scenarios

Technologies

vllm, sglang, TRT-llm

Responsibilities

Develop and optimize LLM/VLM inference engines; Implement cutting-edge LLM Infra technologies into production; Collaborate with search algorithms to design generational improvements for large search systems

Seniority

Senior, hands-on IC

Sourced via tencent · Listed on CareerPlan, which tracks 846,000+ jobs from 20+ sources.