CareerPlanSign in

豆包大模型算法工程师(火山方舟)-RL Infra

北京💼 Full-time🗓 2026-09-28

Core

Improving RL training systems and optimizing performance/stability for large model post-training tasks (Reasoning, Agent, VLM).

Role type

Senior IC large model RL infrastructure engineer

Builds

RL training systems and post-training pipelines for large models

Domain

Large Language Models / Reinforcement Learning / Distributed Systems

Deliverable

production ML models

Required skills

Reinforcement Learning algorithms, Distributed training frameworks, Python, C/C++/Rust/Golang/Java, Data structures and algorithms, High-performance inference engines

Preferred skills

OpenRLHF, VeRL, PyTorch FSDP2, DeviceMesh, DTensor, vLLM, SGLang, Continuous Batching, PagedAttention, Prefix Caching, SRFT, OnPolicy Distillation

Responsibilities

Extend RL Trainer functionality based on Ray, explore Rollout/sampling strategies, integrate Reward systems, manage trajectories in complex Agent Loop tasks, optimize SFT/RL training performance and stability, explore frontier algorithms like Off-Policy RL and DPO/PPO/GRPO

Seniority

Senior, hands-on IC

Sourced via bytedance · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.