CareerPlanGet AI match score →

Research, Vision Expertise

San Francisco💼 Full-time💰 $350,000–$350,000🗓 2026-05-04 → 2026-07-31

Core

Advancing the science of visual perception and multimodal learning by designing architectures that fuse pixels and text, building datasets, and developing representations for real-world comprehension.

Role type

Research, Vision Expertise (Multimodal AI)

Builds

Frontier multimodal models and evaluation tools for visual understanding and reasoning

Domain

Artificial Intelligence / Multimodal Learning / Computer Vision

Deliverable

production ML models

Required skills

Design and analyze large-scale experiments, machine learning fundamentals, distributed compute environments, Python, deep learning frameworks (PyTorch/TensorFlow/JAX), theoretical and empirical grounding

Preferred skills

Visual reasoning and spatial understanding, multimodal architecture design, evaluation frameworks for multimodal tasks, publications in vision-language modeling, probability and statistics

Technologies

PyTorch, TensorFlow, JAX

Responsibilities

Own research projects on training and performance analysis of multimodal AI models, curate and build large-scale datasets and evaluation benchmarks, collaborate with data infrastructure and product teams to create frontier models, publish and present research to advance the community

Seniority

Individual Contributor (Research & Engineering)

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗