语音合成算法工程师(J94365)
Core
Leading R&D for generative speech synthesis, deploying zero-shot/few-shot and cross-lingual TTS technologies into products like digital human live streaming and smart assistants.
Role type
Senior IC machine-learning engineer (speech synthesis) with team management responsibilities
Builds
Production TTS systems for digital human live streaming, smart assistants, in-vehicle, and IoT scenarios
Domain
Artificial Intelligence / Speech Technology
Deliverable
production ML models
Required skills
Generative models (VAE, Flow, Diffusion, LLM, GAN), Zero-shot/few-shot TTS, Cross-lingual/cross-voice synthesis, End-to-end TTS (FastSpeech, Glow-TTS, VITS), Large-scale audio data processing (millions of hours), System architecture from 0-1
Preferred skills
PhD in relevant fields, Deep understanding of data cleaning, augmentation, and simulation pipelines
Technologies
VAE, Flow-based models, Diffusion, LLM, GAN, FastSpeech, Glow-TTS, VITS
Responsibilities
Lead team R&D for generative speech synthesis models; Drive deployment of advanced TTS technologies; Build voice data production and evaluation systems; Collaborate with product teams for scenario scaling; Manage team, talent development, and梯队建设