Research Scientist, Video Foundation Models
Core
Developing next-generation native video and omni foundation models for multimodal generation and understanding, spanning pre-training, continued training, and post-training.
Role type
Senior IC research scientist (video foundation models)
Builds
Large-scale video and multimodal foundation models for high-quality, controllable, and consistent generation
Domain
Generative AI, video generation, multimodal models
Deliverable
production ML models
Required skills
generative modeling, diffusion models, flow matching, DiTs, video generation, multimodal models, world models, large-scale model training, distributed training systems, data curation, model architecture design, post-training, reward modeling, model acceleration, inference systems
Preferred skills
research publications, open-source contributions, independent ownership of ambiguous research problems
Technologies
modern deep learning frameworks, distributed training systems
Responsibilities
Research, develop, and scale native video and multimodal foundation models from early prototypes through large-scale training; Explore new model architectures, training objectives, and conditioning mechanisms; Build and improve large-scale data curation, distributed training, evaluation, and post-training pipelines; Design systematic experiments to understand model scaling, generation quality, controllability, consistency, robustness, and inference efficiency; Collaborate with researchers, engineers, and product teams to shape the technical roadmap; Contribute to research publications and open-source releases
Seniority
Senior, hands-on IC