CareerPlanSign in

光子 AI-高级研究员-语音合成与多模态大模型

Shenzhen, China💼 Full-time🗓 2026-09-28

Core

Research and develop state-of-the-art speech synthesis and generation algorithms (TTS, voice conversion, sound/music generation) using LLMs and multimodal/full-modal large models.

Role type

Senior Research Scientist (Speech AI & Multimodal LLMs)

Builds

Online speech and audio synthesis systems, real-time voice interaction capabilities

Domain

AI, Speech Processing, Multimodal Large Models

Deliverable

production ML models

Required skills

PhD in CS/EE/Signal Processing, Generative Models (Diffusion, Flow Matching, Autoregressive), LLM extension to audio modality, Python, PyTorch, Distributed Training, Top-tier Conference Publications

Preferred skills

Full-duplex/Streaming voice dialogue systems, Real-time interaction modeling

Technologies

PyTorch, Diffusion Models, Flow Matching, Autoregressive Models, LLMs

Responsibilities

Develop and optimize speech/audio synthesis systems for online applications, Explore and advance full-duplex/streaming multimodal LLMs for voice understanding and generation, Collaborate with research and engineering teams to move projects from prototype to production

Seniority

Senior, hands-on IC

Sourced via tencent · Listed on CareerPlan, which tracks 846,000+ jobs from 20+ sources.