Senior Machine Learning Engineer - Speech to Text
Required skills
Deep expertise in speech and audio ML systems (5+ years), including experience with ASR, diarization, VAD, or multilingual speech models. Strong understanding of real-time or streaming ML systems, with experience optimizing for latency and reliability. Experience with transformer-based or hybrid speech architectures (e.g., wav2vec2, Whisper-like systems, conformers, etc.). Solid knowledge of audio preprocessing, feature extraction, and robustness techniques. Strong Python skills and experience deploying ML systems in production. Ability to design and evaluate experiments with rigorous metrics tailored to speech systems (WER, DER, latency, robustness benchmarks). Strong product mindset: you understand that improvements in speech quality directly impact clinician experience. Fluent in English.
Preferred skills
Experience with multilingual speech systems and language detection. Experience in healthcare or other high-reliability domains. Experience working on edge or resource-constrained inference environments.
Technologies
Speech recognition, speaker diarization, voice activity detection (VAD), principal speaker detection, language detection, custom vocabulary, transformer-based or hybrid speech architectures (e.g., wav2vec2, Whisper-like systems, conformers), audio preprocessing, feature extraction, streaming ML systems, Python.
Responsibilities
Contribute to the improvements of core components of our speech stack, including speech recognition, speaker diarization, voice activity detection (VAD), principal speaker detection, language detection and custom vocabulary. Design and improve state-of-the-art speech systems, adapting modern architectures to real-world clinical environments. Improve real-time performance, helping optimize streaming pipelines to balance latency, accuracy, and cost in live clinical workflows. Increase robustness and reliability, ensuring consistent performance across accents, specialties, acoustic conditions, and healthcare organizations. Strengthen our evaluation standards, defining and refining metrics around WER, DER, latency, multilingual performance, and production monitoring. Contribute to the speech data strategy, improving dataset quality, annotation processes, and continuous learning loops from production feedback. Collaborate closely with product, ML, and infrastructure engineers to ensure speech systems integrate seamlessly into downstream note generation and agentic workflows.
Seniority
Senior
Domain
Healthcare