强化学习系统平台工程师 - Seed Model
Core
Building distributed online reinforcement learning system platforms and optimizing performance for O1/O3 chain-of-thought models to explore AGI pathways.
Role type
Senior IC reinforcement learning systems engineer
Builds
Distributed training reward evaluation systems for Agents, Function Calls, and Sandbox environments; Agent frameworks supporting complex interaction RL training; observability and interpretability systems for RL tasks.
Domain
Artificial Intelligence / Reinforcement Learning / Distributed Systems
Deliverable
production ML models
Required skills
Linux environment development, Go/Python/Shell programming, Kubernetes architecture, Ray framework development, distributed system design and maintenance, logical analysis and abstraction
Preferred skills
PyTorch/Megatron-LM/DeepSpeed frameworks, RLHF frameworks (OpenRLHF/VeRL/ChatLearn), sandbox/virtual machine/security container experience, publications in OSDI/SOSP/NSDI/ATC/EuroSys
Technologies
Go, Python, Shell, Kubernetes, Ray, PyTorch, Megatron-LM, DeepSpeed, OpenRLHF, VeRL, ChatLearn
Responsibilities
Construct and optimize distributed online RL system platforms for O1/O3 models; Build distributed training reward evaluation systems for Agent and Function Call scenarios; Develop Agent frameworks supporting complex interaction RL training; Build observability and interpretability systems for RL tasks; Optimize RL task performance to improve model iteration efficiency.
