CareerPlanSign in

Senior Researcher, Speech Synthesis and Multimodal LLM

Singapore-CapitaSky💼 Full-time🗓 2026-07-17 → 2026-09-26

Core

Research and develop advanced speech synthesis, voice conversion, and sound/music generation algorithms using LLMs and multimodal models for online applications.

Role type

Senior Researcher, Speech Synthesis and Multimodal LLM

Builds

Speech and audio synthesis systems for online applications

Domain

Gaming, Speech/Audio Processing, Generative AI

Deliverable

production ML models

Required skills

Speech/audio processing, modern generative models (diffusion, flow matching, autoregressive, codec-based), LLM extension to speech/audio modalities, full-duplex/streaming spoken dialogue systems, Python, deep learning frameworks (PyTorch), distributed training

Preferred skills

Real-time interaction modeling

Technologies

PyTorch

Responsibilities

Research and develop advanced speech synthesis and generation algorithms based on LLMs and multimodal/omnimodal LLMs; Develop and optimize speech and audio synthesis systems for online applications; Explore and advance full-duplex/streaming multimodal LLM capabilities in speech understanding, generation, and real-time spoken interaction; Collaborate cross-functionally with research and engineering teams from prototyping to production

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.