CareerPlanSign in

Applied Scientist, VLM / Vision Language

United States💼 Full-time💰 $200,000–$200,000🗓 2026-05-18 → 2026-09-14

Core

Advancing vision-language models (VLMs) that power intelligent agents in complex, real-world environments through multimodal reasoning and agentic decision-making.

Role type

Senior Applied Scientist (Multimodal AI / VLMs)

Builds

Large-scale multimodal foundation models and intelligent agent workflows

Domain

Artificial Intelligence / Computer Vision / Natural Language Processing

Deliverable

production ML models

Required skills

Multimodal model design, training, and post-training; Cross-modal alignment; SFT, RLHF, DPO, and reward modelling; Evaluation framework design for reasoning and grounding; Large multimodal dataset curation (synthetic and proprietary); Research implementation and experiment design

Preferred skills

Experience with VLMs or multimodal foundation models; Hands-on post-training and alignment expertise

Technologies

SFT, RLHF, DPO, reward modelling frameworks

Responsibilities

Train and fine-tune large-scale vision-language models; Improve multimodal alignment between image and text representations; Apply post-training techniques to strengthen reasoning and factual consistency; Design evaluation frameworks for reasoning quality and robustness; Work with large multimodal datasets including synthetic and proprietary data

Seniority

Senior, hands-on IC with research impact

Sourced via techire · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.