CareerPlanSign in

混元VLM 预训练数据算法工程师(北京/深圳/上海)

Beijing, China💼 Full-time🗓 2026-09-28

Core

Design and implement end-to-end pipelines for collecting, cleaning, and annotating multimodal data to pretrain Vision-Language Models (VLMs), focusing on alignment, robustness, and training efficiency.

Role type

Senior IC multimodal data algorithm engineer

Builds

High-throughput data processing pipelines and curated multimodal datasets for VLM pretraining

Domain

Artificial Intelligence / Multimodal Large Models

Deliverable

production ML models

Required skills

Deep learning, Computer Vision, Natural Language Processing, Transformer architecture, Multimodal alignment, Python, PyTorch, HuggingFace, Spark, Flink, DDP, FSDP, DeepSpeed

Preferred skills

Automatic annotation, Data synthesis, Visual Grounding, LLM-assisted generation, Large-scale dataset construction (100M+)

Sourced via tencent · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.