跨模态算法工程师(J92111)
Core
Research and optimize multimodal large models for image and video understanding to support Baidu's Wenxin base model training and inference capabilities.
Role type
Senior IC multimodal algorithm engineer (large language models)
Builds
Production multimodal large models and inference systems
Domain
AI / Multimodal Large Language Models
Deliverable
production ML models
Required skills
Python, PyTorch, multimodal model architecture, large-scale pretraining, RLHF, data synthesis and augmentation, model fine-tuning
Preferred skills
Experience with LLaVA, Qwen, ERNIE 4.5, top-tier conference publications (ACL/CVPR/ICLR/NeurIPS), open-source project leadership
Technologies
LLaVA, Qwen, ERNIE 4.5
Responsibilities
Optimize base multimodal models for image and video understanding; Conduct RLHF research to explore reasoning limits; Innovate multimodal synthetic data generation; Collaborate with business teams to deploy models; Track and implement cutting-edge international technologies
Seniority
Senior, hands-on IC