上海-多模态算法工程师(J100767)
Core
Research and engineering implementation of multimodal large models (vision-language, audio-language, cross-modal generation) to enhance understanding and generation capabilities.
Role type
Multimodal algorithm engineer (research & engineering)
Builds
Multimodal large models for understanding and generation tasks
Domain
Artificial Intelligence / Multimodal Learning
Deliverable
production ML models
Required skills
Deep learning fundamentals, Transformer architectures, Vision Transformers (ViT), CLIP, Python, PyTorch, PaddlePaddle
Preferred skills
Multimodal learning, Computer Vision, NLP research, Top-tier conference publications (CVPR, ECCV, ACL)
Sourced via baidu · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.