CareerPlanSign in

多模态算法工程师-文档智能

杭州💼 Full-time🗓 2026-09-28

Core

Research and develop multimodal large models for document intelligence, including visual rich document parsing, key information extraction, video summarization, and visual document translation.

Role type

Research-level multimodal algorithm engineer (document intelligence)

Builds

Efficient, scalable multimodal large model architectures for document processing tasks

Domain

Computer Vision + Multimodal AI + Document Intelligence

Deliverable

production ML models

Required skills

C++, Python, computer vision algorithms, multimodal algorithms, machine learning algorithms, academic research

Preferred skills

Published papers in computer vision, patent applications

Technologies

Large language models, multimodal architectures

Responsibilities

Explore applications of multimodal understanding and generation models in document intelligence, follow up on frontier technologies, conduct deep research on key technical challenges

Seniority

Senior, research & hands-on IC

Sourced via bytedance · Listed on CareerPlan, which tracks 855,000+ jobs from 20+ sources.