Senior Machine Learning Engineer
Core
Design and optimize real-time speech systems including recognition, diarization, and voice activity detection for noisy, multilingual, and telehealth environments.
Role type
Senior IC machine learning engineer (speech/audio)
Builds
Production speech systems integrated with note-generation and agentic workflows
Domain
Healthcare technology / Speech and audio machine learning
Deliverable
production ML models
Required skills
ASR, speaker diarization, VAD, multilingual speech models, transformer architectures (wav2vec2, Whisper, Conformer), audio preprocessing, feature extraction, Python, production deployment, experiment design, WER/DER metrics
Preferred skills
Multilingual speech systems, language detection, healthcare domain experience, edge inference, resource-constrained optimization
Technologies
wav2vec2, Whisper, Conformer, Python
Responsibilities
Improve core speech stack components for noisy rooms, overlapping speech, and telehealth calls; Optimize real-time streaming pipelines for latency, accuracy, cost, and reliability; Define and refine speech evaluation metrics including WER, DER, and latency; Improve speech datasets, annotation processes, and continuous learning loops; Collaborate with product and infrastructure engineers to integrate speech systems with note-generation workflows
Seniority
Senior, hands-on IC