Staff / Principal Machine Learning Engineer, Serving - Switzerland
Core
Building and optimizing high-performance, low-latency serving infrastructure for state-of-the-art realtime voice models used in consumer AI applications.
Role type
Staff / Principal Machine Learning Engineer (Serving)
Builds
Production-grade inference systems for voice AI models
Domain
AI / Machine Learning / Real-time Voice Systems
Deliverable
production ML models
Required skills
Inference optimization (vLLM, TRT-LLM), Model acceleration (quantization, distillation, continuous batching, paged attention, speculative decoding), High-performance systems programming (C++, CUDA, Rust, optimized Python), Distributed systems (Kubernetes, Ray, multi-GPU/multi-node inference), Full-cycle model ownership
Preferred skills
Open-source contributions to inference engines, Non-trivial systems programming projects, PhD in CS/Physics/Math or equivalent practical experience
Technologies
vLLM, TRT-LLM, C++, CUDA, Rust, Python, Kubernetes, Ray, NVIDIA GPUs
Responsibilities
Optimize model inference latency and throughput, Profile and tune code for maximum GPU performance, Design and implement distributed serving systems, Containerize and deploy models to production, Ensure system reliability under high concurrency
Seniority
Staff / Principal, hands-on IC with strategic impact