Backend / ML-Ops Engineer — Speech Model Deployment & Inference Optimization
Core
Own infrastructure and pipelines for integrating trained ASR/TTS/Speech-LLM models into production, focusing on scalable serving, GPU optimization, and inference reliability.
Role type
Senior IC ML-Ops Engineer (Speech Model Deployment)
Builds
Production speech inference pipelines for AI voice agents and productivity copilots in healthcare
Domain
Healthcare AI, Speech Recognition, Multimodal Foundation Models
Deliverable
production ML models
Required skills
Triton Inference Server, TensorRT, Kubernetes, GPU scheduling, Python, CI/CD pipelines, observability (Prometheus/Grafana), model versioning (MLflow/DVC), edge inference, mixed precision/quantization
Preferred skills
Streaming ASR/TTS pipelines, DevSecOps, PHI-safe environments, cloud services (AWS/GCP/Azure)
Technologies
Triton, TensorRT, Docker, Kubernetes, Prometheus, Grafana, ELK, MLflow, DVC
Responsibilities
Containerize and deploy speech models with TensorRT/FP16 optimizations; Configure autoscaling on Kubernetes GPU pools; Build health and observability dashboards for latency and WER drift; Implement on-device or edge inference paths; Optimize GPU/CPU utilization and memory footprint for concurrent workloads
Seniority
Senior, hands-on IC