Tech Lead — ASR / TTS / Speech LLM (IC + Mentor)
Core
Lead end-to-end technical development of speech models (ASR, TTS, Speech-LLM) for healthcare applications, guiding a small team on training, data generation, and deployment.
Role type
Senior IC Tech Lead (Speech AI)
Builds
Production speech models and inference systems for patient triage, remote monitoring, and nursing workflow automation.
Domain
Healthcare AI / Speech Processing
Deliverable
production ML models
Required skills
Python, PyTorch, Hugging Face, NeMo, ESPnet, audio data processing, ASR fine-tuning, LoRA/QLoRA, distributed training, LLM alignment, evaluation frameworks (WER/sWER, Entity-F1), inference serving (vLLM/Triton), multi-GPU systems
Preferred skills
Streaming ASR (Transducer/zipformer), context biasing (WFST), bias-aware training, telephony speech characteristics, MLflow logging, quantization, kv-cache optimization
Technologies
PyTorch, Hugging Face, NeMo, ESPnet, vLLM, Triton, MLflow, Livekit, Pipecat, Dify, Kaldi, K2, SpeechBrain
Responsibilities
Prepare and maintain synthetic and real training datasets; fine-tune speech models using CTC/RNN-T or adapter-based recipes; implement evaluation pipelines for clinical applications; optimize inference latency and cost; collaborate with backend and DevOps teams to integrate models; assist in maintaining benchmarks versus commercial baselines.
Seniority
Senior, hands-on IC with mentorship