CareerPlanSign in

微信搜索-LLM大模型推理研发高级工程师-研发优化

Beijing, China💼 Full-time🗓 2026-09-28

Core

Develop and optimize online inference models (SLM, LLM) for WeChat search scenarios to ensure engine performance, stability, and cost-efficiency under high C-end traffic.

Role type

Senior IC LLM inference engineering specialist

Builds

High-performance inference engines for generative and discriminative large models

Domain

Search & Recommendation (Search, Ads, Recommendation)

Deliverable

production ML models

Required skills

C/C++, Linux development, distributed systems principles, LLM algorithm architecture, VLLM, SGLang

Preferred skills

Inference engineering optimization in search/ads/recommendation scenarios

Technologies

VLLM, SGLang

Responsibilities

Develop and optimize online inference models for various sizes; Apply distributed system principles to solve engineering challenges in LLM inference; Implement cutting-edge LLM Infra technologies; Collaborate with search algorithms to design generational model improvements

Seniority

Senior, hands-on IC

Sourced via tencent · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.