CareerPlanSign in

混元多模态强化学习(RL)算法研究员(北京/上海)

Shenzhen, China💼 Full-time🗓 2026-09-28

Core

Researching reinforcement learning algorithms for multimodal models, including diffusion models for image/video generation and autoregressive models for multimodal understanding.

Role type

Senior IC research scientist (multimodal RL)

Builds

Multimodal generative models (image/video) and understanding frameworks

Domain

AI/ML, Multimodal Learning, Reinforcement Learning

Deliverable

production ML models

Required skills

Reinforcement learning algorithms, Diffusion models, Autoregressive models, Deep learning system implementation, Distributed training, GPU/CPU optimization

Preferred skills

Text-to-image/video generation, ACM/NOIP competition experience

Technologies

Diffusion models, Autoregressive models, Distributed training frameworks

Responsibilities

Design and develop RL training frameworks and reward modeling strategies; Optimize large-scale training stability and address reward hacking; Explore next-generation RL paradigms for efficient environment feedback learning.

Sourced via tencent · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.