大模型工程研发实习生(J100640)
Core
Build reward and environment services for reinforcement learning training of large language models.
Role type
Intern, large model engineering (RL training infrastructure)
Builds
RL training reward and environment services for large language models
Domain
Artificial Intelligence / Large Language Models / Reinforcement Learning
Deliverable
production ML models
Required skills
Java, Python, Go, C/C++, PyTorch, PaddlePaddle, Linux, virtualization, containerization, high-performance computing, cloud storage
Preferred skills
large model training, fine-tuning, inference, MDP, Policy, Value, PPO
Responsibilities
Build reward and environment services for RL training; ensure system delivery, high availability, stability, and scalability; solve complex system issues through technical deep dives; conduct technical research and sharing to improve team capabilities.
