Machine Learning Engineer
Core
Design and build production-grade deep learning models for speech/audio and computer vision domains, owning the end-to-end lifecycle from data curation to deployment.
Role type
Senior Machine Learning Engineer (Speech/Audio & Computer Vision)
Builds
Low-latency, scalable AI features for speech/audio (ASR, TTS, noise suppression) and computer vision (detection, segmentation, video understanding)
Domain
Speech/Audio and Computer Vision
Deliverable
production ML models
Required skills
Deep learning model development, dataset curation, architecture design, training pipeline management, model optimization (quantization, pruning), API/microservices design, evaluation benchmarking, distributed training, cloud ML platforms, containerized deployment, system design for ML services
Preferred skills
JAX, knowledge distillation, specific speech frameworks (Wav2Vec 2.0, Whisper, ESPnet, SpeechBrain), specific CV models (YOLO, DETR, ViT, EfficientNet, SAM), inference servers (TensorRT, ONNX Runtime, Triton)
Technologies
PyTorch, TensorFlow, Keras, OpenCV, torchvision, Albumentations, Docker, Kubernetes, AWS SageMaker, GCP Vertex AI, Azure ML, SQL
Responsibilities
Research, design, train, and productionize deep learning models for speech/audio and computer vision use cases; Architect training pipelines capable of handling large-scale datasets; Select and adapt state-of-the-art architectures and fine-tune or distill pre-trained models; Optimize models for inference targeting latency, throughput, and memory budgets; Work with software engineers to integrate models into production systems; Define and track evaluation benchmarks and monitor model performance in production
Seniority
Senior, hands-on IC