多模态大模型视频理解算法研发工程师(J89977)
Core
Research and develop state-of-the-art multimodal large models for video understanding, focusing on tasks like Video QA, captioning, action localization, and event detection.
Role type
Senior IC multimodal large model video understanding algorithm engineer
Builds
Production video understanding models and datasets for product applications
Domain
Artificial Intelligence, Computer Vision, Multimodal Learning
Deliverable
production ML models
Required skills
Transformer, ViT, CNN, RNN, Python, deep learning frameworks, video understanding, multimodal learning, LLM/VLM fine-tuning, distributed training, large-scale data processing
Preferred skills
Publications in top AI conferences (CVPR, ICCV, NeurIPS, etc.)
Technologies
Linux, Git
Responsibilities
Research and implement SOTA video understanding models; build large-scale multimodal datasets; perform distributed training and optimization; collaborate with business teams to deploy algorithms in products.
Seniority
Senior, hands-on IC