Multimodal ML Engineer
Core
Train and ship multimodal models (vision, audio, video, speech) for an AI safety platform that enforces natural-language policies on AI systems.
Role type
Senior IC multimodal machine learning engineer
Builds
Multimodal models (vision-language, audio, speech) and alignment pipelines for an AI safety platform
Domain
AI Safety, Multimodal Machine Learning
Deliverable
production ML models
Required skills
Large-scale deep learning model training, PyTorch with distributed training, Multimodal architecture design, RLHF/alignment (GRPO, DPO, reward modeling), Video/audio sequence modeling, Production model optimization, Multimodal dataset curation, MoE architectures
Preferred skills
Audio signal processing fundamentals
Technologies
PyTorch, DeepSpeed, FSDP, LLaVA, Qwen-VL, InternVL, Whisper, HuBERT, Conformer
Responsibilities
Train and fine-tune large-scale multimodal models from scratch and pretrained checkpoints; Design and run experiments on architecture changes and training recipes; Build and maintain multimodal data pipelines including synthetic data generation; Optimize MoE architectures for efficient inference; Build alignment pipelines across modalities; Optimize models for production latency and serving; Deploy models end-to-end from research to production; Define evaluation metrics for visual QA, spatial reasoning, and audio understanding
Seniority
Senior, hands-on IC