视觉先锋-多模态理解&生成大模型算法实习生计划(J86894)
Core
Research and development of frontier models for multimodal understanding, video/image generation, digital humans, and 3D generation.
Role type
Research intern (multimodal AI)
Builds
Production ML models for content understanding and generation
Domain
Artificial Intelligence, Computer Vision, Natural Language Processing
Deliverable
production ML models
Required skills
Multimodal learning, Deep learning, Computer vision, Natural language processing, Model implementation
Preferred skills
Publications at top conferences (CVPR/ICCV/ECCV/AAAI/ACL/EMNLP/NeurIPS), High ranking in coding/academic competitions (ACM/Kaggle)
Technologies
Deep learning frameworks, Multimodal architectures
Responsibilities
Research and implement state-of-the-art algorithms, Reproduce and improve existing algorithms, Collaborate with product and engineering teams to define technical solutions