混元预训练数据工程负责人
Core
Building and managing large-scale data pipelines and infrastructure to support pre-training data supply for foundation models.
Role type
Senior IC data engineering lead (pre-training)
Builds
Large-scale data pipelines and data asset systems for model pre-training
Domain
AI/ML infrastructure, large-scale data processing
Deliverable
production ML models
Required skills
Large-scale data pipeline architecture, distributed big data processing (Spark/Ray/OSS), data lifecycle management, data quality governance, team leadership, cross-functional collaboration
Preferred skills
Large model pre-training workflows, multimodal data handling
Technologies
Spark, Ray, OSS
Responsibilities
Design and build standardized large-scale data pipelines; Manage the full lifecycle of pre-training data including standards, quality, and versioning; Lead technical breakthroughs in data processing and resolve engineering bottlenecks; Collaborate with model training algorithms and underlying architecture teams.
Seniority
Senior, hands-on IC with team leadership