微信搜索-Agent算法专家
Core
Optimizing WeChat Search Agent capabilities for complex real-world tasks, including DeepSearch/DeepResearch and Agentic workflows.
Role type
Senior IC reinforcement learning engineer (Agentic RL)
Builds
Next-generation large model + Agent + Search technology and product paradigms
Domain
AI / Large Language Models / Search
Deliverable
production ML models
Required skills
PPO, GRPO, DPO, RLHF, CoT, self-reflection, tool learning, distributed training frameworks
Preferred skills
NeurIPS/ICLR/ICML first-author publications on RL or Agents
Technologies
Mid-Train, SFT, GRM, PRM, RLVR, Agentic RL, Context management, Memory
Sourced via tencent · Listed on CareerPlan, which tracks 848,000+ jobs from 20+ sources.
