Researcher, Multimodal Safety
Core
Define and advance multimodal safety research for text, vision, and audio to ensure frontier models behave safely across Chat and multimodal experiences.
Role type
Senior IC multimodal safety researcher (vision-language models)
Builds
Training and evaluation methods for VLMs, including post-training, safety evals, and interventions
Domain
AI safety, multimodal systems, vision-language models
Deliverable
production ML models
Required skills
multimodal model development, vision-language models, modality fusion, image encoders, multimodal post-training, safety evaluations, SFT, RL, data curation, synthetic data, error analysis, cross-modal reasoning, scaling, inference optimization
Preferred skills
video understanding, image generation, audio systems, multimodal reasoning, end-to-end system architecture
Technologies
VLMs, SFT, RL, synthetic data pipelines
Responsibilities
Define and advance multimodal safety research connecting perception to safe behavior; Build training and evaluation methods for VLMs; Collaborate with product and model teams to translate research into safer multimodal experiences
Seniority
Senior, hands-on IC