Applied Scientist, VLM / Vision Language
Core
Advancing vision-language models (VLMs) that power intelligent agents in complex, real-world environments through multimodal reasoning and agentic decision-making.
Role type
Senior Applied Scientist (Multimodal AI / VLMs)
Builds
Large-scale multimodal foundation models and intelligent agent workflows
Domain
Artificial Intelligence / Computer Vision / Natural Language Processing
Deliverable
production ML models
Required skills
Multimodal model design, training, and post-training; Cross-modal alignment; SFT, RLHF, DPO, and reward modelling; Evaluation framework design for reasoning and grounding; Large multimodal dataset curation (synthetic and proprietary); Research implementation and experiment design
Preferred skills
Experience with VLMs or multimodal foundation models; Hands-on post-training and alignment expertise
Technologies
SFT, RLHF, DPO, reward modelling frameworks
Responsibilities
Train and fine-tune large-scale vision-language models; Improve multimodal alignment between image and text representations; Apply post-training techniques to strengthen reasoning and factual consistency; Design evaluation frameworks for reasoning quality and robustness; Work with large multimodal datasets including synthetic and proprietary data
Seniority
Senior, hands-on IC with research impact