CareerPlanSign in

腾讯云-MaaS产品强化学习专家工程师

Shenzhen, China💼 Full-time🗓 2026-09-28

Core

Building product capabilities for the MaaS platform's model fine-tuning system, focusing on RLHF/RLAIF/DPO/GRPO/PPO method selection, iterative improvement, and reasoning enhancement via verifiable rewards and long chain-of-thought optimization.

Role type

Senior IC machine-learning engineer (reinforcement learning for LLMs)

Builds

MaaS platform fine-tuning infrastructure and reasoning-optimized models

Domain

Cloud computing + Large Language Models + Reinforcement Learning

Deliverable

production ML models

Required skills

Reinforcement learning algorithms (RLHF, RLAIF, DPO, GRPO, PPO), Large-scale distributed training, LLM post-training, Reward modeling, Asynchronous training, Tool calling & multi-agent collaboration

Preferred skills

None stated

Technologies

None explicitly named

Responsibilities

Selecting and iterating RL methods for model fine-tuning; Implementing verifiable reward-driven RL training and long chain-of-thought optimization; Building RL infrastructure including large-scale rollout and reward model services; Exploring Agentic RL for platform scenarios; Collaborating on training-evaluation-inference closed loops

Seniority

Senior, hands-on IC

Sourced via tencent · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.