CareerPlanSign in

Senior Machine Learning Engineer, LLM Inference Optimization

Switzerland💼 Full-time🗓 2026-09-24 → 2026-09-25

Core

Drive optimization of large language and vision-language model inference from model artifacts through production deployment to improve latency, throughput, memory efficiency, GPU utilization, reliability, and cost per token.

Role type

Senior IC machine learning engineer (LLM inference optimization)

Builds

High-throughput AI inference serving systems for production environments

Domain

AI Infrastructure / Large Language Models / GPU Computing

Deliverable

production ML models

Required skills

Python, PyTorch, LLM/VLM inference optimization, modern inference engines (vLLM, SGLang, TensorRT-LLM, Triton), transformer inference bottlenecks analysis, quantitative performance trade-off reasoning, complex performance problem diagnosis

Preferred skills

Model compression workflows (quantization, distillation), advanced inference techniques (speculative decoding, KV-cache optimization), agentic workload support, CUDA/Triton, open-source contributions

Technologies

vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, PyTorch, CUDA

Responsibilities

Own optimization initiatives for specific model families and inference serving backends; Evaluate inference engines and recommend serving configurations; Diagnose and resolve model quality, performance, and reliability regressions; Optimize LLM/VLM endpoints for latency, throughput, and cost; Deploy and benchmark modern inference engines; Build and productionize model-compression workflows; Implement advanced inference techniques like speculative decoding and continuous batching; Develop reproducible benchmark harnesses; Partner with GPU kernel and platform engineers to identify bottlenecks; Produce design documentation and performance reports; Contribute to safe production rollouts

Seniority

Senior, hands-on IC

Sourced via lever · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.