Senior Machine Learning Engineer - Speech to Text
Core
Building the intelligence layer for a real-time clinical AI assistant that processes audio from medical encounters to streamline documentation and improve patient care.
Role type
Senior IC machine learning engineer (speech and audio)
Builds
Real-time clinical audio stack including speech recognition, speaker diarization, and voice activity detection for a healthcare AI assistant.
Domain
Healthcare / Speech and Audio Machine Learning
Deliverable
production ML models
Required skills
Speech and audio ML systems (ASR, diarization, VAD), transformer-based speech architectures, real-time/streaming ML optimization, audio preprocessing and feature extraction, Python, production ML deployment, experimental design with speech metrics (WER, DER, latency)
Preferred skills
Multilingual speech systems, healthcare domain experience, edge or resource-constrained inference
Technologies
wav2vec2, Whisper-like systems, conformers
Responsibilities
Contribute to core components of the speech stack (speech recognition, speaker diarization, VAD, language detection), design and improve state-of-the-art speech systems for real-world clinical environments, optimize streaming pipelines for latency and cost, strengthen evaluation standards and metrics, contribute to the speech data strategy and continuous learning loops, collaborate with product and infrastructure engineers on downstream integration
Seniority
Senior, hands-on IC