多模态内容理解算法专家-抖音直播(深圳/北京)
Core
Develop industry-leading multimodal content understanding large models for real-time interaction in Douyin live streaming.
Role type
Senior IC multimodal AI algorithm engineer
Builds
Multimodal dialogue systems, interactive content generation, and captioning features for live streaming
Domain
Live streaming, multimodal AI, large language models (LLM), vision-language models (VLM)
Deliverable
production ML models
Required skills
Multimodal large models (VLM), large language models (LLM), reinforcement learning (RL), model training, fine-tuning, model compression/distillation, data preprocessing, feature extraction, Megatron, Deepspeed
Preferred skills
Visual Chain-of-Thought (CoT) research, real-time multimodal streaming computation, top-tier conference publications, competition awards
Technologies
CNN, VLM, Megatron, Deepspeed
Responsibilities
Train, fine-tune, evaluate, and deploy multimodal models for downstream business applications; optimize data preprocessing, cleaning, annotation, and feature fusion methods for massive multimodal datasets.
Seniority
Senior, hands-on IC