CareerPlanSign in

混元Agent后训练算法工程师(北京/上海/深圳)

Beijing, China💼 Full-time🗓 2026-09-28

Core

Design and implement post-training algorithms for large language model agents, focusing on task planning, memory mechanisms, tool use, and multi-turn dialogue consistency.

Role type

Senior IC machine-learning engineer (LLM agent post-training)

Builds

Post-training pipelines for large-scale data parallel processing and distributed training

Domain

Artificial Intelligence / Large Language Models

Deliverable

production ML models

Required skills

Reinforcement learning, Natural language processing, Transformer architecture, Python, PyTorch, Distributed training (DDP/FSDP/DeepSpeed), Instruction tuning, Reward model training

Preferred skills

Multi-agent collaboration, Hierarchical task planning, Hallucination mitigation, Tool learning

Responsibilities

Design and implement post-training schemes including SFT, RM training, and RLHF/RLAIF; Build post-training data systems for collection, cleaning, and annotation; Optimize agent capabilities in complex task decomposition and cross-domain knowledge transfer; Engineer high-efficiency post-training pipelines supporting distributed training; Track and apply frontier technologies like LLM+Planning and Multi-Agent Interaction.

Sourced via tencent · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.