腾讯云-MaaS产品强化学习专家工程师
Core
Building product capabilities for the MaaS platform's model fine-tuning system, focusing on RLHF/RLAIF/DPO/GRPO/PPO method selection, iterative improvement, and reasoning enhancement via verifiable rewards and long chain-of-thought optimization.
Role type
Senior IC machine-learning engineer (reinforcement learning for LLMs)
Builds
MaaS platform fine-tuning infrastructure and reasoning-optimized models
Domain
Cloud computing + Large Language Models + Reinforcement Learning
Deliverable
production ML models
Required skills
Reinforcement learning algorithms (RLHF, RLAIF, DPO, GRPO, PPO), Large-scale distributed training, LLM post-training, Reward modeling, Asynchronous training, Tool calling & multi-agent collaboration
Preferred skills
None stated
Technologies
None explicitly named
Responsibilities
Selecting and iterating RL methods for model fine-tuning; Implementing verifiable reward-driven RL training and long chain-of-thought optimization; Building RL infrastructure including large-scale rollout and reward model services; Exploring Agentic RL for platform scenarios; Collaborating on training-evaluation-inference closed loops
Seniority
Senior, hands-on IC