CareerPlanSign in

微信小程序-大模型infra工程师-Coding方向

Guangzhou, China💼 Full-time🗓 2026-09-28

Core

Building and optimizing distributed training infrastructure and LLM inference services for WeChat Mini Program Coding Agent scenarios.

Role type

Senior IC machine-learning infrastructure engineer (LLM training & inference)

Builds

Distributed training clusters, high-performance LLM inference services for coding agents

Domain

AI/ML infrastructure, Large Language Models, Cloud Computing

Deliverable

production ML models

Required skills

C/C++/Python, distributed training frameworks (DeepSpeed/Megatron-LM/FSDP), LLM inference frameworks (vLLM/SGLang), memory management, resource scheduling, algorithm design

Preferred skills

Asynchronous reinforcement learning framework development, online stability governance, speculative decoding

Technologies

vLLM, SGLang, DeepSpeed, Megatron-LM, FSDP, C++, Python

Responsibilities

Build and optimize distributed training infrastructure including memory management and communication efficiency; Develop and optimize LLM inference services covering KV Cache management and speculative decoding; Manage GPU resources through fine-grained allocation and flexible scheduling

Sourced via tencent · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.