Senior Researcher, Speech Synthesis and Multimodal LLM
Core
Research and develop advanced speech synthesis, voice conversion, and sound/music generation algorithms using LLMs and multimodal models for online applications.
Role type
Senior Researcher, Speech Synthesis and Multimodal LLM
Builds
Speech and audio synthesis systems for online applications
Domain
Gaming, Speech/Audio Processing, Generative AI
Deliverable
production ML models
Required skills
Speech/audio processing, modern generative models (diffusion, flow matching, autoregressive, codec-based), LLM extension to speech/audio modalities, full-duplex/streaming spoken dialogue systems, Python, deep learning frameworks (PyTorch), distributed training
Preferred skills
Real-time interaction modeling
Technologies
PyTorch
Responsibilities
Research and develop advanced speech synthesis and generation algorithms based on LLMs and multimodal/omnimodal LLMs; Develop and optimize speech and audio synthesis systems for online applications; Explore and advance full-duplex/streaming multimodal LLM capabilities in speech understanding, generation, and real-time spoken interaction; Collaborate cross-functionally with research and engineering teams from prototyping to production
Seniority
Senior, hands-on IC