混元多模态强化学习(RL)算法研究员(北京/上海)
Core
Researching reinforcement learning algorithms for multimodal models, including diffusion models for image/video generation and autoregressive models for multimodal understanding.
Role type
Senior IC research scientist (multimodal RL)
Builds
Multimodal generative models (image/video) and understanding frameworks
Domain
AI/ML, Multimodal Learning, Reinforcement Learning
Deliverable
production ML models
Required skills
Reinforcement learning algorithms, Diffusion models, Autoregressive models, Deep learning system implementation, Distributed training, GPU/CPU optimization
Preferred skills
Text-to-image/video generation, ACM/NOIP competition experience
Technologies
Diffusion models, Autoregressive models, Distributed training frameworks
Responsibilities
Design and develop RL training frameworks and reward modeling strategies; Optimize large-scale training stability and address reward hacking; Explore next-generation RL paradigms for efficient environment feedback learning.