Research Internship Reinforcement Learning (Summer)
Core
Conducting cutting-edge research in reinforcement learning (RL) and large language models (LLMs), focusing on self-distillation, verifiable rewards, and handling extremely large rollouts.
Role type
Research Intern (Reinforcement Learning & LLMs)
Builds
Theoretical mathematical models and practical implementations for LLM training and deployment.
Domain
Artificial Intelligence, Machine Learning, Large Language Models
Deliverable
research
Required skills
Reinforcement learning, Deep learning, Python, PyTorch, TensorFlow, LLM training paradigms, Code generation, Unit testing, Compiler tools
Preferred skills
RLVR, Self-distillation, Large-scale ML experiments
Technologies
PyTorch, TensorFlow
Responsibilities
Conduct literature reviews and implement state-of-the-art algorithms in RL and self-distillation; Design and execute experiments to evaluate proposed methods on code generation and agentic tasks; Develop and maintain codebases for theoretical modeling and practical implementations; Collaborate with researchers to analyze results, refine methodologies, and prepare findings for publication; Contribute to the design of mechanisms for handling large rollouts, such as summarization and hierarchical sub-agents; Document progress, methodologies, and outcomes clearly and comprehensively.