微信搜索-AI Infra 工程师-大模型推理方向 (深圳)(广州)
Core
Design and optimize inference engines for Large Language Models (LLM) and Vision-Language Models (VLM) to support AI Search and intelligent Agent applications.
Role type
Senior IC LLM Infrastructure Engineer (Inference)
Builds
Inference infrastructure for large-scale AI Search and Agent systems
Domain
AI / Large Language Models / Search Technology
Deliverable
production ML models
Required skills
LLM/VLM model architecture, inference engine optimization (vllm/sglang/TRT-llm), operator fusion, quantization strategies, dynamic batching, distributed KV cache optimization, large-scale system design
Preferred skills
Experience with AI hardware configuration, model structure optimization for real-world scenarios
Technologies
vllm, sglang, TRT-llm
Responsibilities
Develop and optimize LLM/VLM inference engines; Implement cutting-edge LLM Infra technologies into production; Collaborate with search algorithms to design generational improvements for large search systems
Seniority
Senior, hands-on IC