Member of Technical Staff, Data & ML Infrastructure for Video Models
Core
Build and scale data pipelines for large video generation models, handling ingestion, parsing, filtering, and dataset curation to prepare high-quality training samples.
Role type
Senior IC data engineer specializing in ML infrastructure for video models
Builds
Scalable data pipelines, annotation workflows, and preprocessing infrastructure for video generation models
Domain
Generative AI, Computer Vision, Video Processing
Deliverable
production ML models
Required skills
Python, AWS (S3, DynamoDB), Kubernetes, PyTorch, distributed data processing, data pipeline orchestration, annotation workflow design, model inference optimization
Preferred skills
Experience with generative video, computer vision, or multimodal ML; training/evaluating smaller supporting models for filtering or quality assessment
Technologies
AWS S3, AWS DynamoDB, Kubernetes, PyTorch, MTurk, Prolific, RunPod
Responsibilities
Build and maintain data pipelines for large video generation models; Design and run annotation workflows across platforms; Train and evaluate smaller supporting models for data filtering; Partner with research teams to scale experimental workflows; Own data quality and identify pipeline bottlenecks; Build internal tools for dataset preparation and monitoring; Drive large-scale pipeline projects from start to finish; Profile and optimize research model inference scripts for preprocessing
Seniority
Senior, hands-on IC