语音对话大模型算法实习生(J106084)
Core
Research and development of end-to-end speech dialogue large models, covering Speech-to-Text, Speech-to-Speech, streaming speech understanding, and generation.
Role type
Research intern (Speech LLM)
Builds
Interactive digital human scenarios
Domain
Artificial Intelligence / Speech Technology
Deliverable
production ML models
Required skills
Python, PyTorch, Transformer architecture, LLM training, ASR, TTS, audio codecs, SFT, reinforcement learning, English technical reading
Preferred skills
Experience in dialogue systems, role-playing, voice cloning, natural interruption management
Technologies
PyTorch, Transformer, SFT, RLHF
Responsibilities
Participate in model training, fine-tuning, inference optimization, and evaluation; Build and clean speech dialogue datasets; Explore post-training methods to improve instruction following and role consistency; Establish evaluation systems for content quality, naturalness, latency, and robustness; Reproduce state-of-the-art research in Speech LLM.
Seniority
Intern