语音算法实习生(J84146)
Core
Research and develop large-scale speech synthesis and understanding systems, specifically LLM-based TTS and end-to-end Speech2Speech architectures.
Role type
Research intern (Speech AI / Large Language Models)
Builds
LLM-based Text-to-Speech (TTS) and Speech2Speech systems
Domain
Artificial Intelligence / Speech Processing / Large Language Models
Deliverable
production ML models
Required skills
Python, C/C++, PyTorch, TTS algorithms, ASR algorithms, VITS, VALLE, FishSpeech, CosyVoice, LLM theory
Preferred skills
End-to-end speech modeling, multimodal alignment, model evaluation and optimization
Responsibilities
Decouple and analyze modules in latest speech synthesis and understanding solutions; Participate in R&D of speech Encodec, Decoder, and multimodal alignment modules; Evaluate and optimize speech large models
Sourced via baidu · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.