CareerPlanSign in

语音合成算法工程师(J94365)

北京市💼 Full-time🗓 2026-07-21 → 2026-09-28

Core

Leading R&D for generative speech synthesis, deploying zero-shot/few-shot and cross-lingual TTS technologies into products like digital human live streaming and smart assistants.

Role type

Senior IC machine-learning engineer (speech synthesis) with team management responsibilities

Builds

Production TTS systems for digital human live streaming, smart assistants, in-vehicle, and IoT scenarios

Domain

Artificial Intelligence / Speech Technology

Deliverable

production ML models

Required skills

Generative models (VAE, Flow, Diffusion, LLM, GAN), Zero-shot/few-shot TTS, Cross-lingual/cross-voice synthesis, End-to-end TTS (FastSpeech, Glow-TTS, VITS), Large-scale audio data processing (millions of hours), System architecture from 0-1

Preferred skills

PhD in relevant fields, Deep understanding of data cleaning, augmentation, and simulation pipelines

Technologies

VAE, Flow-based models, Diffusion, LLM, GAN, FastSpeech, Glow-TTS, VITS

Responsibilities

Lead team R&D for generative speech synthesis models; Drive deployment of advanced TTS technologies; Build voice data production and evaluation systems; Collaborate with product teams for scenario scaling; Manage team, talent development, and梯队建设

Sourced via baidu · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.