Rewrite
## About the Role
We are looking for a passionate and highly skilled Machine Learning Engineer – Voice AI to lead the development of our multilingual and multi-dialect speech systems. The ideal candidate should have hands-on experience in traditional speech processing, Text-to-Speech (TTS), Automatic Speech Recognition (ASR), audio signal processing, and Voice AI technologies.
This role is critical to our long-term vision of building and self-hosting production-grade Voice AI solutions tailored for multiple industry use cases across India. Since we work extensively with regional Indian languages and dialects, we are looking for someone who can take ownership of the Voice AI domain with strong technical leadership and a research-driven mindset.
## Key Responsibilities
- Design, develop, and optimize Voice AI systems using traditional and modern speech processing techniques.
- Build multilingual and multi-dialect speech solutions for Indian regional languages.
- Develop and improve Text-to-Speech (TTS), speech enhancement, pronunciation modeling, and voice adaptation pipelines.
- Work on Automatic Speech Recognition (ASR), speaker identification, speaker verification, and keyword spotting systems.
- Design robust audio preprocessing, feature extraction, and post-processing pipelines.
- Improve speech quality, intelligibility, naturalness, and dialect adaptation.
- Work with traditional speech processing techniques, digital signal processing (DSP), and statistical speech models where applicable.
- Build scalable training and inference pipelines for self-hosted Voice AI systems.
- Optimize low-latency inference for production deployments.
- Collaborate with product, engineering, and data teams to deploy production-ready Voice AI solutions.
- Evaluate models using objective speech quality metrics and human evaluation.
- Stay up to date with advancements in speech processing, Voice AI, and audio machine learning.
## Required Skills & Qualifications
- 2+ years of experience in Machine Learning, Speech Processing, Voice AI, or related domains.
- Strong understanding of traditional speech processing techniques, including:
- Digital Signal Processing (DSP)
- MFCC, LPC, PLP, Mel Spectrograms
- Fourier Transform (FFT), Short-Time Fourier Transform (STFT), and Filter Banks
- Speech segmentation and Voice Activity Detection (VAD)
- Experience with traditional Text-to-Speech systems, statistical parametric speech synthesis, concatenative synthesis, or modern neural TTS architectures.
- Good understanding of speech feature extraction, spectrograms, vocoders, and audio codecs.
- Experience with Automatic Speech Recognition (ASR) pipelines and acoustic modeling is preferred.
- Strong Python programming skills.
- Hands-on experience with PyTorch, TensorFlow, or similar machine learning frameworks.
- Experience handling multilingual speech datasets and audio corpora.
- Familiarity with Linux, Docker, GPU training, and inference optimization.
- Understanding of speech evaluation metrics such as WER, CER, MOS, PESQ, and STOI.
- Ability to independently own and drive Voice AI research and production initiatives.
## Preferred Qualifications
- Experience working with Indian language datasets and dialect adaptation.
- Experience with speaker recognition, speaker diarization, or voice biometrics.
- Familiarity with Kaldi, ESPnet, HTK, CMU Sphinx, OpenSMILE, Praat, or similar speech processing toolkits.
- Experience building self-hosted Voice AI infrastructure.
- Knowledge of conversational AI, telephony systems, and speech analytics.
- Research contributions, open-source projects, or published work in speech processing or speech synthesis.
## What We Offer
- Opportunity to work on cutting-edge multilingual Voice AI systems.
- Access to large proprietary speech datasets.
- Freedom to experiment, research, and build production-grade Voice AI infrastructure.
- High-impact role with ownership and technical leadership opportunities.
- Collaborative and innovation-driven work environment.
## Ideal Candidate
We are looking for someone who is deeply interested in speech processing and Voice AI, comfortable taking ownership of the complete speech pipeline—from audio preprocessing and feature engineering to speech synthesis, recognition, and deployment. The ideal candidate is excited about solving challenges in Indian languages and dialects, capable of leading Voice AI initiatives end-to-end, and committed to building scalable, production-ready speech systems using both traditional speech processing techniques and modern machine learning approaches.
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.