混元VLM 预训练数据算法工程师(北京/深圳/上海)
Core
Design and implement end-to-end pipelines for collecting, cleaning, and annotating multimodal data to pretrain Vision-Language Models (VLMs), focusing on alignment, robustness, and training efficiency.
Role type
Senior IC multimodal data algorithm engineer
Builds
High-throughput data processing pipelines and curated multimodal datasets for VLM pretraining
Domain
Artificial Intelligence / Multimodal Large Models
Deliverable
production ML models
Required skills
Deep learning, Computer Vision, Natural Language Processing, Transformer architecture, Multimodal alignment, Python, PyTorch, HuggingFace, Spark, Flink, DDP, FSDP, DeepSpeed
Preferred skills
Automatic annotation, Data synthesis, Visual Grounding, LLM-assisted generation, Large-scale dataset construction (100M+)
Sourced via tencent · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.