CareerPlanSign in

Senior ML Engineer - Kimchi (LLM Inference Optimization

UK🌐 Remote💼 Full-time🗓 2026-05-26 → 2026-08-07

Core

Optimizing LLM inference throughput, latency, and KV cache utilization on cloud infrastructure to improve customer performance and margins.

Role type

Senior IC machine-learning engineer (LLM inference optimization)

Builds

Autonomous decision-making layer for Kubernetes and cloud environments that matches workloads to cost-efficient LLM configurations

Domain

Cloud-native AI infrastructure, Kubernetes, Large Language Model serving

Deliverable

production ML models

Required skills

Python (production services), vLLM, SGLang, TensorRT-LLM, quantization tradeoffs, distributed systems (collective communication, sharding), measurement-driven optimization, kernel-level tuning

Preferred skills

None stated

Technologies

vLLM, SGLang, TensorRT-LLM, PyTorch, CUDA, Kubernetes, gRPC, ClickHouse, PostgreSQL, GCP Pub/Sub, AWS, GCP, Azure, GitLab CI, ArgoCD, Prometheus, Grafana, Loki, Tempo

Responsibilities

Tune continuous batching, speculative decoding, and chunked prefill to maximize GPU throughput; Profile and fix latency bottlenecks (TTFT, TPOT); Manage KV cache via paged attention and prefix caching; Quantize weights and activations without quality regression; Optimize cold starts and memory footprint; Implement distributed inference topologies; Define technical direction and benchmarking strategies

Seniority

Senior, hands-on IC

Sourced via adzuna · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.