多模态数据工程师 - Seed Model
Core
Manage the full lifecycle of trillion-scale multimodal data (video, image) to support large model pre-training, SFT fine-tuning, and evaluation for products like Doubao and Jimeng.
Role type
Senior Multimodal Data Engineer (Seed Model)
Builds
High-concurrency, low-latency distributed multimodal data processing engines and data insight platforms.
Domain
AI / Large Language Models / Multimodal Data Engineering
Deliverable
production ML models
Required skills
Python or Golang, Distributed system architecture, Data modeling, Big data frameworks (Hadoop, Spark, Flink, Ray), Data governance, OCR/ASR integration, Transformer/CNN architecture understanding
Preferred skills
LLM/VLM SFT methodologies, Inference acceleration, Video semantic segmentation, Top-tier academic publications, End-to-end data pipeline construction
Technologies
Hive, ClickHouse, MySQL, MongoDB, ElasticSearch, Hadoop, Spark, Flink, Ray, Transformer, CNN
Responsibilities
Design storage architecture and distributed processing for massive multimodal data; Optimize data cleaning, structuring, and annotation workflows; Build data insight platforms to analyze data distribution and quality impact on model training; Collaborate with algorithm teams to close the data-training-evaluation loop; Productize data capabilities into tools for the team.