微信小程序-大模型infra工程师-Coding方向
Core
Building and optimizing distributed training infrastructure and LLM inference services for WeChat Mini Program Coding Agent scenarios.
Role type
Senior IC machine-learning infrastructure engineer (LLM training & inference)
Builds
Distributed training clusters, high-performance LLM inference services for coding agents
Domain
AI/ML infrastructure, Large Language Models, Cloud Computing
Deliverable
production ML models
Required skills
C/C++/Python, distributed training frameworks (DeepSpeed/Megatron-LM/FSDP), LLM inference frameworks (vLLM/SGLang), memory management, resource scheduling, algorithm design
Preferred skills
Asynchronous reinforcement learning framework development, online stability governance, speculative decoding
Technologies
vLLM, SGLang, DeepSpeed, Megatron-LM, FSDP, C++, Python
Responsibilities
Build and optimize distributed training infrastructure including memory management and communication efficiency; Develop and optimize LLM inference services covering KV Cache management and speculative decoding; Manage GPU resources through fine-grained allocation and flexible scheduling