Machine Learning Engineer (Singapore)
Core
Build and scale distributed data pipelines for ingesting, processing, and delivering large-scale video and multimodal data for model training.
Role type
Senior IC machine learning engineer (data infrastructure & curation)
Builds
Production data pipelines, dataset curation systems, and VLM-based captioning workflows for social AI models
Domain
Social AI, video processing, multimodal data engineering
Deliverable
production ML models
Required skills
Python, distributed data processing (PySpark, Ray), orchestration (Airflow), containerization (Docker, Kubernetes), cloud storage (AWS, GCS, Azure), VLM captioning, CLIP-based filtering, video processing (FFmpeg, PyAV, DALI, OpenCV)
Preferred skills
Experience with near-deduplication pipelines, aesthetic scoring models, dataset versioning strategies
Technologies
PySpark, Ray, Airflow, Docker, Kubernetes, AWS, GCS, Azure, FFmpeg, PyAV, DALI, OpenCV, Decord, torchvision, PyTorchVideo, torchaudio
Responsibilities
Design and scale distributed data pipelines for preprocessing and dataset generation; Own workflow orchestration, job scheduling, monitoring, and failure recovery; Implement containerized pipeline infrastructure; Optimize cloud-based data storage and movement; Define best practices for dataset storage layout and versioning; Build curation pipelines for video and image content selection; Develop VLM-based captioning and metadata generation workflows; Apply quality and aesthetic scoring models for data selection; Build tooling for deduplication workflows; Analyze dataset composition and iterate on curation logic
Seniority
Senior, hands-on IC