混元多模态强化学习后训练算法工程师(框架协同方向)(北京/上海)
Core
Designing and optimizing post-training algorithms (RLHF, DPO, SFT) for multimodal large models, acting as a technical bridge between algorithm and framework teams.
Role type
Senior IC multimodal reinforcement learning post-training algorithm engineer (framework collaboration)
Builds
Multimodal large language models and post-training pipelines
Domain
Artificial Intelligence / Multimodal Large Models / Reinforcement Learning
Deliverable
production ML models
Required skills
Python, PyTorch, Transformer architectures, Diffusion models, RLHF/DPO/SFT algorithms, training stability optimization, reward function design, root cause analysis, technical documentation
Preferred skills
Experience with Megatron-LM, DeepSpeed, VLLM, VERL, OpenRLHF, cross-modal alignment, hardware optimization collaboration
Technologies
PyTorch, Megatron-LM, DeepSpeed, VLLM, VERL, OpenRLHF
Responsibilities
Translate post-training algorithm principles into functional requirements for framework architecture; Lead the setup, optimization, and evaluation of post-training workflows; Conduct research on frontier techniques and resolve training bottlenecks; Collaborate with framework and hardware teams to ensure technical solution implementation.
Seniority
Senior, hands-on IC