CareerPlanSign in

多模态算法工程师-抖音内容理解

上海💼 Full-time🗓 2026-09-28

Core

Building and optimizing foundational multimodal large models for content understanding and generation across video, audio, image, and text domains to power search, recommendation, and advertising.

Role type

Senior IC multimodal algorithm engineer (large language models & generative AI)

Builds

Systematic multimodal foundation models and generative model capabilities for Douyin's ecosystem

Domain

Social media, video streaming, and content recommendation

Deliverable

production ML models

Required skills

deep learning, large language models, multimodal models, generative models, mathematical foundations, data engineering, model training, training/inference framework iteration, model evaluation metrics

Preferred skills

experience with short video and image-text algorithms, publications in top-tier conferences (NIPS, ICML, CVPR, etc.), competition experience

Technologies

PyTorch, TensorFlow, Hugging Face, distributed training systems, vector databases

Responsibilities

Research and develop open-set content understanding models; Train and maintain multimodal foundation models; Iterate on training and inference frameworks; Explore and integrate latest industry technologies

Seniority

Senior, hands-on IC

Sourced via bytedance · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.