视觉多模态应用算法工程师/专家(视频通话方向) - Seed Model
Core
Optimizing post-training for video call models to enhance multi-turn dialogue, visual perception, search, memory, and agent capabilities, while developing new features like active response and full-duplex audio/video.
Role type
Senior IC multimodal algorithm engineer (video calls)
Builds
Video call models and intelligent hardware capabilities for consumer apps and enterprise clients
Domain
AI, Multimodal Large Language Models, Video Technology
Deliverable
production ML models
Required skills
Video large model training, Visual perception, Multi-turn dialogue optimization, Agent invocation, End-to-end experience optimization, Active response features, Full-duplex audio/video, Research paper publication in top conferences
Preferred skills
Experience with intelligent hardware, Master's degree in AI/CS/Automation/Mathematics
Technologies
Video large models, Agent frameworks, Search integration
Responsibilities
Optimize post-training for video call models including visual perception and memory, Develop new video call features like active response and full-duplex, Optimize intelligent hardware capabilities for end-to-end user experience, Explore and apply frontier innovative technologies to application effects