Principal Scientist - Data Pipeline Engineer
Core
Architect and scale multimodal data processing pipelines and infrastructure to turn billions of raw assets into training-ready data for Adobe Firefly's foundation models.
Role type
Principal Scientist - Data Pipeline Engineer (Senior IC)
Builds
Distributed, GPU-accelerated systems for data ingestion, processing, and delivery for multimodal foundation models.
Domain
Generative AI, Multimodal Models (Image, Video, Audio), Large-Scale Data Infrastructure
Deliverable
production ML models
Required skills
Distributed systems architecture, GPU inference optimization, Data curation for training, Systems-level programming (C++, Rust, Go, Java), Database and storage systems at scale, Python, Ray/Spark frameworks
Preferred skills
Experience with VLMs and LLMs, Hands-on technical leadership, Full-stack systems knowledge
Technologies
Ray, Spark, Python, C++, Rust, Go, Java, GPU clusters, Distributed storage, Large-scale databases
Responsibilities
Architect and optimize large-scale distributed pipelines processing billions of assets, Scale inference throughput and eliminate bottlenecks, Design systems for storing and serving billions of data points, Partner with modeling teams to translate training needs into pipeline requirements, Own architecture decisions for database, storage, and compute utilization, Operate as a hands-on technical leader bridging data engineering and applied ML
Seniority
Principal, hands-on IC with broad technical influence