CareerPlanSign in

Multimodal ML Engineer

Paris💼 Full-time🗓 2026-07-02 → 2026-09-26

Core

Train and ship multimodal models (vision, audio, video, speech) for an AI safety platform that enforces natural-language policies on AI systems.

Role type

Senior IC multimodal machine learning engineer

Builds

Multimodal models (vision-language, audio, speech) and alignment pipelines for an AI safety platform

Domain

AI Safety, Multimodal Machine Learning

Deliverable

production ML models

Required skills

Large-scale deep learning model training, PyTorch with distributed training, Multimodal architecture design, RLHF/alignment (GRPO, DPO, reward modeling), Video/audio sequence modeling, Production model optimization, Multimodal dataset curation, MoE architectures

Preferred skills

Audio signal processing fundamentals

Technologies

PyTorch, DeepSpeed, FSDP, LLaVA, Qwen-VL, InternVL, Whisper, HuBERT, Conformer

Responsibilities

Train and fine-tune large-scale multimodal models from scratch and pretrained checkpoints; Design and run experiments on architecture changes and training recipes; Build and maintain multimodal data pipelines including synthetic data generation; Optimize MoE architectures for efficient inference; Build alignment pipelines across modalities; Optimize models for production latency and serving; Deploy models end-to-end from research to production; Define evaluation metrics for visual QA, spatial reasoning, and audio understanding

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.