大模型Code/Agent后训练算法研究员-(深圳)or(北京)or
Core
Researching and optimizing Agentic RL algorithms, data construction, and simulation environments for Code and Agent applications.
Role type
Senior Research Scientist (Agentic RL & Large Model Code/Agent)
Builds
High-quality Code/Agent training datasets, scalable Agent simulation environments, and optimized Agentic RL training infrastructure.
Domain
Artificial Intelligence, Large Language Models, Reinforcement Learning
Deliverable
production ML models
Required skills
Deep learning algorithms, distributed training and inference acceleration, Agentic RL, data construction and governance, simulation environment design
Preferred skills
Large-scale reinforcement learning, large model Code/Agent R&D, top-tier conference publications, leading open-source projects
Technologies
Deep learning frameworks, distributed training systems
Responsibilities
Construct and govern high-quality, diverse Code/Agent training datasets; build and optimize high-availability Agent simulation environments; train and optimize Agentic RL algorithms for Code/Agent scenarios.
Seniority
Senior, hands-on IC