Machine Learning Engineer (Voice AI)
Core
Lead the development of multilingual and multi-dialect speech systems, focusing on traditional speech processing, TTS, and ASR for Indian regional languages.
Role type
Senior IC machine learning engineer (voice AI)
Builds
Production-grade Voice AI solutions for multiple industry use cases
Domain
Voice AI, Speech Processing, Indian Regional Languages
Deliverable
production ML models
Required skills
Digital Signal Processing (DSP), MFCC/LPC/PLP/Mel Spectrograms, FFT/STFT/Filter Banks, Speech segmentation/VAD, Statistical parametric speech synthesis, Neural TTS architectures, Speech feature extraction, Vocoder/Codec knowledge, ASR pipelines, Acoustic modeling, Python, PyTorch/TensorFlow, Multilingual dataset handling, Linux/Docker/GPU optimization, Speech evaluation metrics (WER/CER/MOS/PESQ/STOI)
Preferred skills
Indian language dataset experience, Speaker recognition/diarization/biometrics, Kaldi/ESPnet/HTK/CMU Sphinx/OpenSMILE/Praat, Self-hosted Voice AI infrastructure, Conversational AI/Telephony/Speech analytics, Research contributions
Technologies
PyTorch, TensorFlow, Kaldi, ESPnet, HTK, CMU Sphinx, OpenSMILE, Praat
Responsibilities
Design and optimize Voice AI systems using traditional and modern techniques; Build multilingual speech solutions for Indian languages; Develop TTS, speech enhancement, pronunciation modeling, and voice adaptation pipelines; Work on ASR, speaker identification, verification, and keyword spotting; Design audio preprocessing, feature extraction, and post-processing pipelines; Improve speech quality, intelligibility, and dialect adaptation; Build scalable training and inference pipelines for self-hosted systems; Optimize low-latency inference for production; Evaluate models using objective metrics and human evaluation
Seniority
Senior, hands-on IC