Senior Applied Researcher
Core
Designing and training vision-language models (VLMs) to interpret, reason, and act on large-scale first-person video data in complex real-world environments.
Role type
Senior Applied Researcher (Multimodal AI / VLMs)
Builds
Scalable training and evaluation pipelines for VLMs, post-training approaches (SFT, RLHF), and efficient inference systems.
Domain
Artificial Intelligence, Multimodal Learning, Computer Vision, Natural Language Processing
Deliverable
production ML models
Required skills
Deep learning model training, Transformer architectures, Vision-language model development, Large-scale dataset handling, Model optimization, Video data processing, Long-context modeling, Parameter-efficient tuning
Preferred skills
Experience deploying models into production, Edge and server-side inference design, Benchmarking for spatial and behavioral understanding
Technologies
Transformers, SFT, RLHF, Video datasets
Responsibilities
Designing and training VLMs on large-scale video datasets, Developing post-training approaches including SFT, RLHF, and parameter-efficient tuning, Building scalable training and evaluation pipelines, Exploring long-context and temporal modelling, Designing efficient systems across edge and server-side inference, Defining benchmarks for spatial and behavioural understanding
Seniority
Senior, hands-on IC