CareerPlanSign in

Senior Researcher, Speech Synthesis and Multimodal LLM|高级研究员 - 语音合成与多模态大模型

Tokyo, Japan💼 Full-time🗓 2026-09-28

Core

Research and develop advanced speech synthesis and generation algorithms (e.g., TTS, voice conversion, sound/music generation) based on LLMs and multimodal/omnimodal LLMs for online applications.

Role type

Senior Researcher, Speech Synthesis and Multimodal LLM

Builds

Speech synthesis systems, audio generation models, and real-time spoken interaction capabilities for games and online applications.

Domain

Gaming, Speech/Audio Processing, Generative AI

Deliverable

production ML models

Required skills

LLMs, speech/audio processing, diffusion models, flow matching, autoregressive models, codec-based approaches, speech tokenizers, multimodal adapters, speech-text joint training, Python, PyTorch, distributed training, top-tier academic publications

Preferred skills

full-duplex systems, streaming spoken dialogue systems, real-time interaction modeling

Technologies

PyTorch, LLMs, diffusion, flow matching, autoregressive models, codecs

Responsibilities

Research and develop speech synthesis and generation algorithms; Optimize speech and audio synthesis systems for online applications; Explore full-duplex/streaming multimodal LLM capabilities; Collaborate cross-functionally from prototyping to production.

Sourced via tencent · Listed on CareerPlan, which tracks 878,000+ jobs from 20+ sources.