CareerPlanSign in

Research Scientist, Post-Training — Video Generation

Palo Alto HQ🌐 Remote💼 Full-time🗓 2026-05-16 → 2026-09-26

Core

Refining video generation models using RL alignment and building robust video reward models for real-time creative platforms.

Role type

Staff/Lead Research Scientist (RL Post-Training & Generative Modeling)

Builds

RL-aligned video diffusion/flow-matching models and video reward models

Domain

Generative AI, Video Generation, Reinforcement Learning

Deliverable

production ML models

Required skills

RL post-training, preference optimization, generative modeling, diffusion models, flow-matching models, PyTorch, multi-node distributed training

Preferred skills

video reward model development, VLM-as-judge, large-scale preference data collection, model distillation, video-specific failure mode analysis

Technologies

PyTorch, diffusion models, flow-matching models

Responsibilities

Run RL post-training (preference optimization, online RL) for video models at multi-node scale; Build and validate video reward models; Own post-training evaluation including human preference studies; Distill RL-tuned models to efficient samplers

Seniority

Staff/Lead, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.