微信-AI Infra工程师-大模型训练与RL方向
Core
Design and optimize distributed training frameworks for large-scale LLMs (100B-1T parameters) and build RL infrastructure systems to enable efficient training and inference synergy.
Role type
Senior IC AI Infrastructure Engineer (LLM Training & RL)
Builds
Distributed training frameworks, RL training systems, and inference engines for large language models.
Domain
Artificial Intelligence / Large Language Models / Reinforcement Learning
Deliverable
production ML models
Required skills
Python, system programming, complex system debugging, distributed training principles, Megatron-LM, DeepSpeed, PyTorch FSDP, PPO, GRPO, DPO, vLLM, SGlang, Verl, ROLL, AReal
Preferred skills
None stated
Technologies
Megatron-LM, DeepSpeed, PyTorch FSDP, vLLM, SGlang, Verl, ROLL, AReal
Responsibilities
Design and develop core code for distributed training frameworks based on Megatron-LM/DeepSpeed; Build and optimize RL training frameworks (PPO/GRPO/DPO) addressing resource scheduling and communication bottlenecks; Collaborate with algorithm teams to integrate and scale open-source RL frameworks (Verl, Slime, ROLL, AReal) for business scenarios.