大模型训练框架研发工程师-强化学习/精调/蒸馏方向
Core
Design and develop core modules for large model training frameworks focusing on reinforcement learning, fine-tuning, and knowledge distillation to improve training efficiency and usability.
Role type
Senior IC large model training framework engineer (RL/fine-tuning/distillation)
Builds
Lightweight training frameworks (e.g., LLama-Factory, swift) and distributed training strategies for large models
Domain
Artificial Intelligence / Large Language Models / High-Performance Computing
Deliverable
production ML models
Required skills
Python, C++, distributed training strategies, memory optimization, CUDA programming, operator fusion, collective communication algorithms (NCCL, MPI), RL algorithms (GRPO, PPO)
Preferred skills
Megatron-LM, DeepSpeed, veRL, Slime, AReaL, MoE architecture, FlashMLA, EPLB, DualPipe
Responsibilities
Design and optimize core modules for RL, fine-tuning, and distillation; Optimize distributed training strategies to resolve memory and communication bottlenecks; Develop toolchains for rapid model fine-tuning and multi-hardware adaptation; Translate latest research results into framework features; Collaborate with product teams and write technical documentation.