语音/多模态大模型算法工程师(Speech/Omni/Agent方向) - Data语音
Core
Research and develop end-to-end multi-modal large models (Speech/Omni/Agent) for AIGC, focusing on cross-modal fusion, low-latency interaction, and enterprise application deployment.
Role type
Senior IC multi-modal large model algorithm engineer (Speech/Omni/Agent)
Builds
End-to-end Omni models, Multi-Agent systems, and enterprise-grade AI solutions for smart cockpits, customer service, and productivity tools.
Domain
AI / Large Language Models / Speech & Audio / Multi-modal Systems
Deliverable
production ML models
Required skills
Multi-modal large model development, Speech language models, LLMs, AI Agent system design, Tool use, Complex reasoning, Task planning, Multi-agent collaboration, Reinforcement learning, Model training/inference/deployment, Performance optimization
Preferred skills
Publications in top-tier conferences (NeurIPS, ICML, ICLR, CVPR, ACL, Interspeech, ICASSP), Core contributions to open-source multi-modal/Agent systems, Experience in enterprise productivity scenarios
Technologies
End-to-end model architectures, Multi-agent frameworks, Reinforcement learning algorithms, Speech synthesis and understanding pipelines
