机器学习数据工程师 - Data AML
Core
Designing and developing high-efficiency, high-availability data pipelines for multimodal data production, cleaning, and storage to support recommendation, advertising, and search systems.
Role type
Senior Machine Learning Data Engineer
Builds
Standardized multimodal feature data production, cleaning, and storage pipelines for recommendation and ad systems
Domain
Internet / Machine Learning / Data Engineering
Deliverable
production ML models
Required skills
Spark, Hive, Java, Python, data warehouse implementation, PB-scale cluster optimization, multimodal data processing
Preferred skills
PyTorch, TensorFlow, large model training data characteristics, data annotation systems
Technologies
Spark, Hive, PyTorch, TensorFlow
Responsibilities
Design and develop efficient data pipelines for multimodal data production and storage; Build a cleaning framework for PB-scale data using Spark; Construct multimodal feature caching and storage systems to improve throughput and reduce latency; Participate in building model deployment pipelines and collaborate with algorithm teams for system optimization.