Datascientist
Core
Building and fine-tuning voice-related models from scratch including speech-to-text, speaker diarization, audio classification, and LLM-integrated speech systems.
Role type
Senior IC machine-learning engineer (voice/audio)
Builds
Custom models for speech recognition, transcription, audio segmentation, and speaker identification; robust data pipelines for large audio datasets.
Domain
Voice technology, audio signal processing, machine learning
Deliverable
production ML models
Required skills
Speech recognition, speaker diarization, voice activity detection, model fine-tuning, signal processing, feature extraction, data augmentation, Python, PyTorch, NumPy, Scikit-learn, Hugging Face, end-to-end ML pipelines, cloud platforms (GCP, AWS), containerized environments
Preferred skills
LLM + speech integration, real-time systems, streaming inference, multilingual ASR, dialect modeling, experiment tracking tools (DVC, MLflow, W&B)
Technologies
PyTorch, NumPy, Scikit-learn, Hugging Face, Whisper, wav2vec, HuBERT, Conformer, GCP, AWS
Responsibilities
Lead development of custom models for speech recognition, transcription, audio segmentation, and speaker identification; Build robust data pipelines for collecting, preprocessing, cleaning, and labeling large audio datasets; Fine-tune and evaluate state-of-the-art open-source models on proprietary datasets; Design experiments and benchmark models for quality, latency, and domain adaptability; Work with product teams to embed voice capabilities into real-time applications; Maintain scalable training, evaluation, and inference workflows.
Seniority
Senior, hands-on IC