混元Agent后训练算法工程师(北京/上海/深圳)
Core
Design and implement post-training algorithms for large language model agents, focusing on task planning, memory mechanisms, tool use, and multi-turn dialogue consistency.
Role type
Senior IC machine-learning engineer (LLM agent post-training)
Builds
Post-training pipelines for large-scale data parallel processing and distributed training
Domain
Artificial Intelligence / Large Language Models
Deliverable
production ML models
Required skills
Reinforcement learning, Natural language processing, Transformer architecture, Python, PyTorch, Distributed training (DDP/FSDP/DeepSpeed), Instruction tuning, Reward model training
Preferred skills
Multi-agent collaboration, Hierarchical task planning, Hallucination mitigation, Tool learning
Responsibilities
Design and implement post-training schemes including SFT, RM training, and RLHF/RLAIF; Build post-training data systems for collection, cleaning, and annotation; Optimize agent capabilities in complex task decomposition and cross-domain knowledge transfer; Engineer high-efficiency post-training pipelines supporting distributed training; Track and apply frontier technologies like LLM+Planning and Multi-Agent Interaction.