CareerPlanSign in

强化学习系统平台工程师 - Seed Model

北京💼 Full-time🗓 2026-09-28

Core

Building distributed online reinforcement learning system platforms and optimizing performance for O1/O3 chain-of-thought models to explore AGI pathways.

Role type

Senior IC reinforcement learning systems engineer

Builds

Distributed training reward evaluation systems for Agents, Function Calls, and Sandbox environments; Agent frameworks supporting complex interaction RL training; observability and interpretability systems for RL tasks.

Domain

Artificial Intelligence / Reinforcement Learning / Distributed Systems

Deliverable

production ML models

Required skills

Linux environment development, Go/Python/Shell programming, Kubernetes architecture, Ray framework development, distributed system design and maintenance, logical analysis and abstraction

Preferred skills

PyTorch/Megatron-LM/DeepSpeed frameworks, RLHF frameworks (OpenRLHF/VeRL/ChatLearn), sandbox/virtual machine/security container experience, publications in OSDI/SOSP/NSDI/ATC/EuroSys

Technologies

Go, Python, Shell, Kubernetes, Ray, PyTorch, Megatron-LM, DeepSpeed, OpenRLHF, VeRL, ChatLearn

Responsibilities

Construct and optimize distributed online RL system platforms for O1/O3 models; Build distributed training reward evaluation systems for Agent and Function Call scenarios; Develop Agent frameworks supporting complex interaction RL training; Build observability and interpretability systems for RL tasks; Optimize RL task performance to improve model iteration efficiency.

Sourced via bytedance · Listed on CareerPlan, which tracks 853,000+ jobs from 20+ sources.