多模态算法工程师-抖音内容理解
Core
Building and optimizing foundational multimodal large models for content understanding and generation across video, audio, image, and text domains to power search, recommendation, and advertising.
Role type
Senior IC multimodal algorithm engineer (large language models & generative AI)
Builds
Systematic multimodal foundation models and generative model capabilities for Douyin's ecosystem
Domain
Social media, video streaming, and content recommendation
Deliverable
production ML models
Required skills
deep learning, large language models, multimodal models, generative models, mathematical foundations, data engineering, model training, training/inference framework iteration, model evaluation metrics
Preferred skills
experience with short video and image-text algorithms, publications in top-tier conferences (NIPS, ICML, CVPR, etc.), competition experience
Technologies
PyTorch, TensorFlow, Hugging Face, distributed training systems, vector databases
Responsibilities
Research and develop open-set content understanding models; Train and maintain multimodal foundation models; Iterate on training and inference frameworks; Explore and integrate latest industry technologies
Seniority
Senior, hands-on IC