Senior Researcher, Speech Synthesis and Multimodal LLM|高级研究员 - 语音合成与多模态大模型
Core
Research and develop advanced speech synthesis and generation algorithms (TTS, voice conversion, sound/music generation) based on LLMs and multimodal/omnimodal LLMs for online applications.
Role type
Senior Researcher, Speech Synthesis and Multimodal LLM
Builds
Speech synthesis systems, audio generation models, and real-time spoken interaction capabilities for games.
Domain
Gaming, Speech/Audio Processing, Generative AI
Deliverable
production ML models
Required skills
LLMs, speech/audio processing, diffusion models, flow matching, autoregressive models, codec-based approaches, speech tokenizers, multimodal adapters, speech-text joint training, full-duplex systems, streaming dialogue systems, real-time interaction modeling, Python, PyTorch, distributed training
Preferred skills
None stated
Technologies
PyTorch
Responsibilities
Research and develop speech synthesis algorithms based on LLMs; Optimize speech and audio synthesis systems for online applications; Explore full-duplex/streaming multimodal LLM capabilities; Collaborate cross-functionally from prototyping to production
Seniority
Senior, hands-on IC