CareerPlanSign in

混元大语言模型后训练算法工程师-RM方向(北京/深圳/上海)

Beijing, China💼 Full-time🗓 2026-09-28

Core

Develop and optimize reward systems and RLHF algorithms to enhance large language models in instruction following, reasoning, and value alignment; build personalized user modeling and long-term memory mechanisms.

Role type

Senior IC large language model post-training algorithm engineer (Reward Modeling & Personalization)

Builds

High-quality reward models, personalized LLMs with memory capabilities, and data synthesis pipelines for model iteration

Domain

Artificial Intelligence / Large Language Models / Machine Learning

Deliverable

production ML models

Required skills

Transformer architecture expertise, RLHF and Reward Modeling, Personalized LLM development, Long-term memory/RAG optimization, Python, PyTorch/TensorFlow, Distributed training (Megatron-LM, DeepSpeed, vLLM), Data synthesis (SFT, Self-Instruct), Model evaluation metrics

Preferred skills

Experience with user profiling and recommendation systems, High-impact publications (NeurIPS, ICLR, ICML, ACL, EMNLP), Open source contributions (HuggingFace), Training/optimizing billion-parameter models

Sourced via tencent · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.