混元大语言模型后训练算法工程师-RM方向(北京/深圳/上海)
Core
Develop and optimize reward systems and RLHF algorithms to enhance large language models in instruction following, reasoning, and value alignment; build personalized user modeling and long-term memory mechanisms.
Role type
Senior IC large language model post-training algorithm engineer (Reward Modeling & Personalization)
Builds
High-quality reward models, personalized LLMs with memory capabilities, and data synthesis pipelines for model iteration
Domain
Artificial Intelligence / Large Language Models / Machine Learning
Deliverable
production ML models
Required skills
Transformer architecture expertise, RLHF and Reward Modeling, Personalized LLM development, Long-term memory/RAG optimization, Python, PyTorch/TensorFlow, Distributed training (Megatron-LM, DeepSpeed, vLLM), Data synthesis (SFT, Self-Instruct), Model evaluation metrics
Preferred skills
Experience with user profiling and recommendation systems, High-impact publications (NeurIPS, ICLR, ICML, ACL, EMNLP), Open source contributions (HuggingFace), Training/optimizing billion-parameter models