Machine Learning Engineer
Core
Design and build production-grade deep learning models for speech/audio and computer vision domains, owning the end-to-end lifecycle from data curation to deployment.
Role type
Senior Machine Learning Engineer (Speech/Audio & Computer Vision)
Builds
Low-latency, scalable AI features for speech/audio (ASR, TTS, noise suppression) and computer vision (detection, segmentation, video understanding) use cases.
Domain
Speech processing, audio analysis, computer vision, deep learning
Deliverable
production ML models
Required skills
Deep learning model design and training, dataset curation and preprocessing, model optimization (quantization, pruning, distillation), API and microservices design, evaluation benchmarking, distributed training, cloud ML platform usage, containerized deployment, system design for ML services
Preferred skills
JAX, Wav2Vec 2.0, Whisper, ESPnet, SpeechBrain, torchaudio, librosa, YOLO, DETR, ViT, EfficientNet, SAM, OpenCV, torchvision, Albumentations, TensorRT, ONNX Runtime, TorchScript, DeepSpeed, Triton Inference Server, SQL, CI/CD practices
Technologies
PyTorch, TensorFlow, Keras, AWS SageMaker, GCP Vertex AI, Azure ML, Docker, Kubernetes, ECS, Python
Responsibilities
Research, design, train, and productionize deep learning models for speech/audio and computer vision; Architect training pipelines for large-scale datasets; Select and adapt SOTA architectures and fine-tune pre-trained models; Optimize models for inference latency and throughput; Integrate models into production systems and design serving APIs; Define and track evaluation benchmarks and monitor production performance; Stay current with research literature and implement relevant techniques
Seniority
Senior, hands-on IC