CareerPlanSign in

Machine Learning Engineer

New Delhi, India🌐 Remote💼 Full-time🗓 2026-07-23 → 2026-09-25

Core

Design and build production-grade deep learning models for speech/audio and computer vision domains, owning the end-to-end lifecycle from data curation to deployment.

Role type

Senior Machine Learning Engineer (Speech/Audio & Computer Vision)

Builds

Low-latency, scalable AI features for speech/audio (ASR, TTS, noise suppression) and computer vision (detection, segmentation, video understanding)

Domain

Speech/Audio and Computer Vision

Deliverable

production ML models

Required skills

Deep learning model development, dataset curation, architecture design, training pipeline management, model optimization (quantization, pruning), API/microservices design, evaluation benchmarking, distributed training, cloud ML platforms, containerized deployment, system design for ML services

Preferred skills

JAX, knowledge distillation, specific speech frameworks (Wav2Vec 2.0, Whisper, ESPnet, SpeechBrain), specific CV models (YOLO, DETR, ViT, EfficientNet, SAM), inference servers (TensorRT, ONNX Runtime, Triton)

Technologies

PyTorch, TensorFlow, Keras, OpenCV, torchvision, Albumentations, Docker, Kubernetes, AWS SageMaker, GCP Vertex AI, Azure ML, SQL

Responsibilities

Research, design, train, and productionize deep learning models for speech/audio and computer vision use cases; Architect training pipelines capable of handling large-scale datasets; Select and adapt state-of-the-art architectures and fine-tune or distill pre-trained models; Optimize models for inference targeting latency, throughput, and memory budgets; Work with software engineers to integrate models into production systems; Define and track evaluation benchmarks and monitor model performance in production

Seniority

Senior, hands-on IC

Sourced via pinpoint · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.