原生多模态算法研究员
Core
Research and develop foundational native multimodal models for unified understanding and high-quality generation across image, video, audio, and text.
Role type
Senior IC multimodal algorithm researcher
Builds
Native multimodal base models
Domain
Artificial Intelligence / Multimodal Learning
Deliverable
production ML models
Required skills
Multimodal pretraining, Vision Transformers (ViT), autoregressive models, distributed training frameworks (DeepSpeed, Megatron-LM), CUDA programming, high-performance operator development, Python
Preferred skills
Publications in top-tier conferences (CVPR, ICLR, NeurIPS), experience with ultra-large-scale model training
Technologies
PyTorch, DeepSpeed, Megatron-LM, CUDA
Responsibilities
Design unified modal representation and multi-scale/long-sequence modeling strategies; Optimize model architecture for high-resolution image and long video scenarios; Track and integrate cutting-edge multimodal research advancements.