CareerPlanSign in

Backend / ML-Ops Engineer — Speech Model Deployment & Inference Optimization

Bengaluru💼 Full-time🗓 2025-11-07 → 2026-09-26

Core

Own infrastructure and pipelines for integrating trained ASR/TTS/Speech-LLM models into production, focusing on scalable serving, GPU optimization, and inference reliability.

Role type

Senior IC ML-Ops Engineer (Speech Model Deployment)

Builds

Production speech inference pipelines for AI voice agents and productivity copilots in healthcare

Domain

Healthcare AI, Speech Recognition, Multimodal Foundation Models

Deliverable

production ML models

Required skills

Triton Inference Server, TensorRT, Kubernetes, GPU scheduling, Python, CI/CD pipelines, observability (Prometheus/Grafana), model versioning (MLflow/DVC), edge inference, mixed precision/quantization

Preferred skills

Streaming ASR/TTS pipelines, DevSecOps, PHI-safe environments, cloud services (AWS/GCP/Azure)

Technologies

Triton, TensorRT, Docker, Kubernetes, Prometheus, Grafana, ELK, MLflow, DVC

Responsibilities

Containerize and deploy speech models with TensorRT/FP16 optimizations; Configure autoscaling on Kubernetes GPU pools; Build health and observability dashboards for latency and WER drift; Implement on-device or edge inference paths; Optimize GPU/CPU utilization and memory footprint for concurrent workloads

Seniority

Senior, hands-on IC

Sourced via lever · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.