Senior / Lead Machine Learning Engineer, Serving - Serbia
Core
Building and optimizing high-performance, low-latency serving infrastructure for state-of-the-art realtime voice models used in consumer AI applications.
Role type
Senior IC Machine Learning Engineer (Inference Optimization)
Builds
Realtime voice inference systems, optimized serving frameworks, and scalable APIs for consumer-facing AI products.
Domain
AI/ML, Realtime Inference, Voice Technology
Deliverable
production ML models
Required skills
Inference optimization (vLLM, TRT-LLM), Model acceleration (quantization, distillation, continuous batching, paged attention, speculative decoding), High-performance systems programming (C++, CUDA, Rust, optimized Python), Distributed systems (Kubernetes, Ray, multi-GPU/multi-node inference), Full-cycle model ownership
Preferred skills
Open-source contributions to inference engines, Non-trivial systems programming projects, PhD in CS/Physics/Math
Technologies
vLLM, TRT-LLM, C++, CUDA, Rust, Python, Kubernetes, Ray, NVIDIA GPUs
Responsibilities
Optimize model inference latency and throughput, Profile and squeeze performance from GPU hardware, Manage distributed inference clusters, Containerize and deploy models to production, Design benchmarks for latency and reliability problems
Seniority
Senior, hands-on IC