Summer Camp - Agentic RL /大模型平台策略推理优化实习生(J100476)
Core
Optimizing inference efficiency and training robustness for large language models (LLMs) in real-world agent scenarios, including code execution and multi-turn dialogue.
Role type
Intern, LLM Platform Strategy & Agentic RL Research
Builds
Inference acceleration solutions and Agentic RL training pipelines for the Baidu Qianfan MaaS platform
Domain
Artificial Intelligence / Large Language Models / Reinforcement Learning
Deliverable
production ML models
Required skills
LLM architecture and inference mechanics, vLLM/SGLang framework internals, speculative decoding techniques, Reinforcement Learning theory (PPO, GRPO), LLM training pipelines (SFT, RM, RL), Agent system design (tool use, ReAct, function calling)
Preferred skills
Experience with DeepSeek MTP/Mimo MTP/Eagle3/DFlash implementations, complex multi-step task handling
Technologies
vLLM, SGLang, PPO, GRPO, DeepSeek MTP, Eagle3, DFlash
Responsibilities
Reproduce and improve speculative decoding acceleration schemes; Build Agentic RL training loops for complex business tasks
Seniority
Intern