光子 AI-高级研究员-语音合成与多模态大模型
Core
Research and develop state-of-the-art speech synthesis and generation algorithms (TTS, voice conversion, sound/music generation) using LLMs and multimodal/full-modal large models.
Role type
Senior Research Scientist (Speech AI & Multimodal LLMs)
Builds
Online speech and audio synthesis systems, real-time voice interaction capabilities
Domain
AI, Speech Processing, Multimodal Large Models
Deliverable
production ML models
Required skills
PhD in CS/EE/Signal Processing, Generative Models (Diffusion, Flow Matching, Autoregressive), LLM extension to audio modality, Python, PyTorch, Distributed Training, Top-tier Conference Publications
Preferred skills
Full-duplex/Streaming voice dialogue systems, Real-time interaction modeling
Technologies
PyTorch, Diffusion Models, Flow Matching, Autoregressive Models, LLMs
Responsibilities
Develop and optimize speech/audio synthesis systems for online applications, Explore and advance full-duplex/streaming multimodal LLMs for voice understanding and generation, Collaborate with research and engineering teams to move projects from prototype to production
Seniority
Senior, hands-on IC