Research, Vision Expertise
Core
Advancing the science of visual perception and multimodal learning by designing architectures that fuse pixels and text, building datasets, and developing representations for real-world comprehension.
Role type
Research, Vision Expertise (Multimodal AI)
Builds
Frontier multimodal models and evaluation tools for visual understanding and reasoning
Domain
Artificial Intelligence / Multimodal Learning / Computer Vision
Deliverable
production ML models
Required skills
Design and analyze large-scale experiments, machine learning fundamentals, distributed compute environments, Python, deep learning frameworks (PyTorch/TensorFlow/JAX), theoretical and empirical grounding
Preferred skills
Visual reasoning and spatial understanding, multimodal architecture design, evaluation frameworks for multimodal tasks, publications in vision-language modeling, probability and statistics
Technologies
PyTorch, TensorFlow, JAX
Responsibilities
Own research projects on training and performance analysis of multimodal AI models, curate and build large-scale datasets and evaluation benchmarks, collaborate with data infrastructure and product teams to create frontier models, publish and present research to advance the community
Seniority
Individual Contributor (Research & Engineering)
