深圳-AI Infra 强化学习工程师(基座研发方向)(J101230)
Core
Develop high-scalability decoupled reinforcement learning training frameworks and large-scale post-training platforms to optimize training efficiency and stability for AI models.
Role type
Senior IC reinforcement learning infrastructure engineer (training systems)
Builds
Decoupled RL training frameworks, large-scale post-training platforms, and distributed resource scheduling systems
Domain
AI Infrastructure / Reinforcement Learning Systems
Deliverable
production ML models
Required skills
Python, Go, C++, PyTorch, DeepSpeed, Megatron, distributed training, system scheduling, reinforcement learning algorithms
Preferred skills
veRL, Slime, vLLM, SGLang, K8s, Ray, Reasoning RL, Agentic RL, open source contributions, top-tier conference papers
Technologies
K8s, Ray, Mooncake, vLLM, SGLang, PyTorch, DeepSpeed, Megatron
Responsibilities
Develop decoupled RL training frameworks; optimize asynchronous training paradigms and inference scheduling; build large-scale post-training platforms with fault tolerance; design distributed resource scheduling systems; build observability and automated experimentation platforms
Seniority
Senior, hands-on IC