Research Scientist (Singapore)
Core
Drive foundational research on video generation models, focusing on post-training methods, data pipeline infrastructure, and aligning generation quality with human judgments.
Role type
Senior IC research scientist (video generation & post-training)
Builds
Scalable data pipelines, distillation methods for diffusion/flow models, reward models, and preference-based fine-tuning systems
Domain
Generative AI, video generation, post-training optimization
Deliverable
production ML models
Required skills
Large-scale data systems design, distributed data processing (PySpark/Ray), workflow orchestration (Airflow), container orchestration (Kubernetes), cloud data storage optimization, video/media processing (FFmpeg/PyAV/DALI/OpenCV), post-training research (distillation, RLHF, DPO), Python, PyTorch/JAX
Preferred skills
Publications at top-tier venues (NeurIPS, ICML, ICLR, CVPR, ICCV, ECCV)
Technologies
PySpark, Ray, Airflow, Docker, Kubernetes, AWS, GCS, Azure, FFmpeg, PyAV, DALI, OpenCV, PyTorch, JAX
Responsibilities
Build scalable systems for ingesting and preprocessing large-scale video data; Design and scale distributed data pipelines for dataset generation and refreshes; Own workflow orchestration, job scheduling, and failure recovery; Implement containerized pipeline infrastructure; Optimize cloud-based data storage and movement; Research distillation methods for diffusion and flow-based video generation models; Develop reward models and preference-based fine-tuning pipelines; Analyze base model behavior to inform pretraining decisions
Seniority
Senior, hands-on IC