多模态大模型算法工程师(J97393)
Core
Develop and optimize multimodal large language models for AI content generation (AIGC) across text, image, video, and audio to enhance user experience in products like Baidu Wenku and TeraBox.
Role type
Senior IC multimodal large model algorithm engineer
Builds
AIGC core algorithms and production multimodal models for document and content management products
Domain
AI / Large Language Models / Multimodal Learning
Deliverable
production ML models
Required skills
Natural language processing, computer vision, reinforcement learning, large-scale pretraining, SFT, LoRA, fine-tuning, knowledge distillation, PyTorch, TensorFlow, MXNet, PaddlePaddle
Preferred skills
Publications in top-tier conferences (ACL, EMNLP, NIPS, ICML, ICLR, CVPR, ICCV, ECCV)
Responsibilities
Implement SFT and pretraining for large models; optimize training strategies for modality alignment and reinforcement learning; drive productization of AIGC algorithms; analyze product needs to drive growth through technical innovation