ML Research Scientist I/II, Multimodal Data Extraction
Core
Develop foundation models to autonomously read, interpret, and structure scientific knowledge across text, images, and experimental data in the physical sciences.
Role type
ML Research Scientist (Multimodal Data Extraction)
Builds
AI systems that extract and structure knowledge from diverse scientific sources; scalable pipelines for unstructured and heterogeneous scientific data.
Domain
Physical sciences (materials science, chemistry) + Multimodal AI
Deliverable
production ML models
Required skills
Machine learning, NLP, vision–language modeling, training and fine-tuning LLMs and multimodal models, data structures and representations in physical sciences
Preferred skills
Multimodal fusion architectures, scientific document parsing (OCR, table extraction), knowledge graph construction, handling noisy real-world scientific data
Technologies
PyTorch, Hugging Face Transformers
Responsibilities
Research and develop AI systems for knowledge extraction; Design and fine-tune large language and multi-modal models; Build scalable pipelines for unstructured scientific data; Collaborate with domain experts to align data with discovery workflows; Publish research advancing multimodal understanding.
Seniority
PhD level, research-focused IC