CareerPlanSign in

AI应用技术Leader-集团信息系统

上海💼 Full-time🗓 2026-09-28

Core

Synthesize, clean, and curate pre-training data for large language models; optimize reward models and reinforcement learning algorithms to improve model reasoning and alignment.

Role type

Senior IC machine-learning engineer (RL & LLM alignment)

Builds

Production LLMs with enhanced reasoning capabilities and reward models

Domain

Artificial Intelligence / Large Language Models / Reinforcement Learning

Deliverable

production ML models

Required skills

Data synthesis, data cleaning, reinforcement learning, reward modeling, SFT, algorithm design, Python programming, model architecture understanding

Preferred skills

Reinforcement learning project experience, data-driven work in recommendation systems, rapid prototyping, competitive programming (NOI/ACM)

Responsibilities

Develop data synthesis pipelines to improve pre-training data quality; Optimize reward models and RL algorithms for high-value data feedback loops; Collaborate on SFT stages to address discrimination capability gaps; Organize annotation efforts for reward model training.

Sourced via bytedance · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.