Senior Engineer — ASR / TTS / Speech LLM (Training + Eval + Integration)
Core
Train, evaluate, and integrate speech recognition (ASR), text-to-speech (TTS), and speech LLM models for healthcare voice agents and nursing productivity tools.
Role type
Senior IC machine-learning engineer (speech/audio)
Builds
AI voice agents for patient triage, remote monitoring, and nurse productivity copilots integrated into EHR systems.
Domain
Healthcare technology / Speech AI
Deliverable
production ML models
Required skills
Python, PyTorch, Hugging Face, NeMo, ESPnet, audio data processing, ASR fine-tuning (CTC/RNN-T, Transducer, zipformer), LoRA/adapters, distributed training, evaluation frameworks (WER, Entity-F1, PESQ, STOI), MLflow
Preferred skills
model deployment workflows, bias-aware training, context list conditioning, LM rescoring, WFST boosts, neural re-scoring
Technologies
PyTorch, Hugging Face, NeMo, ESPnet, MLflow, WebRTC, Epic, Athena
Responsibilities
Prepare and maintain synthetic and real training datasets for STT/TTS/LLM models; Build scripts for data selection, augmentation, and corpus curation; Fine-tune models on multi-GPU systems; Implement evaluation pipelines to measure WER, entity F1, and latency; Experiment with bias-aware training and context list conditioning; Collaborate with backend and DevOps teams to integrate trained models into inference stacks; Support creation of context biasing APIs and LM rescoring paths; Assist in maintaining benchmarks versus commercial baselines.
Seniority
Senior, hands-on IC