CareerPlanSign in

微信-AI Infra工程师-大模型训练与RL方向

Beijing, China💼 Full-time🗓 2026-09-28

Core

Design and optimize distributed training frameworks for large-scale LLMs (100B-1T parameters) and build RL infrastructure systems to enable efficient training and inference synergy.

Role type

Senior IC AI Infrastructure Engineer (LLM Training & RL)

Builds

Distributed training frameworks, RL training systems, and inference engines for large language models.

Domain

Artificial Intelligence / Large Language Models / Reinforcement Learning

Deliverable

production ML models

Required skills

Python, system programming, complex system debugging, distributed training principles, Megatron-LM, DeepSpeed, PyTorch FSDP, PPO, GRPO, DPO, vLLM, SGlang, Verl, ROLL, AReal

Preferred skills

None stated

Technologies

Megatron-LM, DeepSpeed, PyTorch FSDP, vLLM, SGlang, Verl, ROLL, AReal

Responsibilities

Design and develop core code for distributed training frameworks based on Megatron-LM/DeepSpeed; Build and optimize RL training frameworks (PPO/GRPO/DPO) addressing resource scheduling and communication bottlenecks; Collaborate with algorithm teams to integrate and scale open-source RL frameworks (Verl, Slime, ROLL, AReal) for business scenarios.

Sourced via tencent · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.