Senior Deep Learning Scientist, Speech Synthesis
Core
Develop high-impact Speech AI product Riva by training speech synthesis models and improving conversational AI experiences for millions of customers.
Role type
Senior Deep Learning Scientist (Speech Synthesis)
Builds
Riva (Speech AI product)
Domain
AI / Speech Technology / Conversational AI
Deliverable
production ML models
Required skills
Python, PyTorch, CNNs, RNNs/LSTMs, Transformers, speech synthesis model training, DSP, feature extraction (FFT, MFCC, Mel spectrograms), Git/Gerrit/GitLab
Preferred skills
C++, CUDA, cuDNN, TensorRT, multilingual TTS, voice cloning, text normalization, inverse text normalization, multilingual G2P systems
Technologies
PyTorch, CUDA, cuDNN, TensorRT, Git, Gerrit, GitLab
Responsibilities
Train Speech Synthesis mel-spectrogram and vocoder models; Measure, benchmark, and analyze model performance, accuracy, and bias; Maintain the TTS model evaluation system; Improve processes for speech data processing, augmentation, filtering, and TTS training set preparation; Build knowledge of TTS datasets for training and evaluation; Collaborate with cross-functional teams on new features, improvements, and issue triage.
Seniority
Senior, hands-on IC