CareerPlanGet AI match score →

Senior ML Engineer - Kimchi (LLM Inference Optimization)

European Union🌐 Remote💼 Full-time🗓 2026-07-10 → 2026-07-31

Core

Optimizing LLM inference throughput, latency, and KV cache utilization on cloud infrastructure to reduce costs and improve performance for customers.

Role type

Senior IC machine-learning engineer (LLM inference optimization)

Builds

Autonomous decision-making layer for matching workloads to cost-efficient LLM configurations and serving settings

Domain

Cloud-native AI infrastructure, Kubernetes, Large Language Model serving

Deliverable

production ML models

Required skills

Python (production services), vLLM/SGLang/TensorRT-LLM, quantization tradeoffs, distributed systems (collective communication, sharding), measurement-driven optimization, kernel-level tuning

Preferred skills

None stated

Technologies

vLLM, SGLang, TensorRT-LLM, PyTorch, CUDA, Kubernetes, gRPC, ClickHouse, PostgreSQL, GCP Pub/Sub, AWS, Azure, GitLab CI, ArgoCD, Prometheus, Grafana, Loki, Tempo

Responsibilities

Tune kernels and schedulers to maximize GPU throughput; profile and fix latency bottlenecks (TTFT, TPOT); implement KV cache optimization strategies (paged attention, prefix caching); perform empirical quantization without quality regression; optimize memory footprint and cold starts; design distributed inference topologies; set technical direction and benchmarking standards

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗