CareerPlanSign in

混元多模态强化学习后训练算法工程师(框架协同方向)(北京/上海)

Shenzhen, China💼 Full-time🗓 2026-09-28

Core

Designing and optimizing post-training algorithms (RLHF, DPO, SFT) for multimodal large models, acting as a technical bridge between algorithm and framework teams.

Role type

Senior IC multimodal reinforcement learning post-training algorithm engineer (framework collaboration)

Builds

Multimodal large language models and post-training pipelines

Domain

Artificial Intelligence / Multimodal Large Models / Reinforcement Learning

Deliverable

production ML models

Required skills

Python, PyTorch, Transformer architectures, Diffusion models, RLHF/DPO/SFT algorithms, training stability optimization, reward function design, root cause analysis, technical documentation

Preferred skills

Experience with Megatron-LM, DeepSpeed, VLLM, VERL, OpenRLHF, cross-modal alignment, hardware optimization collaboration

Technologies

PyTorch, Megatron-LM, DeepSpeed, VLLM, VERL, OpenRLHF

Responsibilities

Translate post-training algorithm principles into functional requirements for framework architecture; Lead the setup, optimization, and evaluation of post-training workflows; Conduct research on frontier techniques and resolve training bottlenecks; Collaborate with framework and hardware teams to ensure technical solution implementation.

Seniority

Senior, hands-on IC

Sourced via tencent · Listed on CareerPlan, which tracks 848,000+ jobs from 20+ sources.