CareerPlanSign in

视觉多模态应用算法工程师/专家(视频通话方向) - Seed Model

北京💼 Full-time🗓 2026-09-28

Core

Optimizing post-training for video call models to enhance multi-turn dialogue, visual perception, search, memory, and agent capabilities, while developing new features like active response and full-duplex audio/video.

Role type

Senior IC multimodal algorithm engineer (video calls)

Builds

Video call models and intelligent hardware capabilities for consumer apps and enterprise clients

Domain

AI, Multimodal Large Language Models, Video Technology

Deliverable

production ML models

Required skills

Video large model training, Visual perception, Multi-turn dialogue optimization, Agent invocation, End-to-end experience optimization, Active response features, Full-duplex audio/video, Research paper publication in top conferences

Preferred skills

Experience with intelligent hardware, Master's degree in AI/CS/Automation/Mathematics

Technologies

Video large models, Agent frameworks, Search integration

Responsibilities

Optimize post-training for video call models including visual perception and memory, Develop new video call features like active response and full-duplex, Optimize intelligent hardware capabilities for end-to-end user experience, Explore and apply frontier innovative technologies to application effects

Sourced via bytedance · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.