CareerPlanSign in

语音/多模态大模型算法工程师(Speech/Omni/Agent方向) - Data语音

上海💼 Full-time🗓 2026-09-28

Core

Research and develop end-to-end multi-modal large models (Speech/Omni/Agent) for AIGC, focusing on cross-modal fusion, low-latency interaction, and enterprise application deployment.

Role type

Senior IC multi-modal large model algorithm engineer (Speech/Omni/Agent)

Builds

End-to-end Omni models, Multi-Agent systems, and enterprise-grade AI solutions for smart cockpits, customer service, and productivity tools.

Domain

AI / Large Language Models / Speech & Audio / Multi-modal Systems

Deliverable

production ML models

Required skills

Multi-modal large model development, Speech language models, LLMs, AI Agent system design, Tool use, Complex reasoning, Task planning, Multi-agent collaboration, Reinforcement learning, Model training/inference/deployment, Performance optimization

Preferred skills

Publications in top-tier conferences (NeurIPS, ICML, ICLR, CVPR, ACL, Interspeech, ICASSP), Core contributions to open-source multi-modal/Agent systems, Experience in enterprise productivity scenarios

Technologies

End-to-end model architectures, Multi-agent frameworks, Reinforcement learning algorithms, Speech synthesis and understanding pipelines

Sourced via bytedance · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.