Senior ML Engineer - Kimchi (LLM Inference Optimization
Core
Optimizing LLM inference throughput, latency, and KV cache utilization on cloud infrastructure to improve customer performance and margins.
Role type
Senior IC machine-learning engineer (LLM inference optimization)
Builds
Autonomous decision-making layer for Kubernetes and cloud environments that matches workloads to cost-efficient LLM configurations
Domain
Cloud-native AI infrastructure, Kubernetes, Large Language Model serving
Deliverable
production ML models
Required skills
Python (production services), vLLM, SGLang, TensorRT-LLM, quantization tradeoffs, distributed systems (collective communication, sharding), measurement-driven optimization, kernel-level tuning
Preferred skills
None stated
Technologies
vLLM, SGLang, TensorRT-LLM, PyTorch, CUDA, Kubernetes, gRPC, ClickHouse, PostgreSQL, GCP Pub/Sub, AWS, GCP, Azure, GitLab CI, ArgoCD, Prometheus, Grafana, Loki, Tempo
Responsibilities
Tune continuous batching, speculative decoding, and chunked prefill to maximize GPU throughput; Profile and fix latency bottlenecks (TTFT, TPOT); Manage KV cache via paged attention and prefix caching; Quantize weights and activations without quality regression; Optimize cold starts and memory footprint; Implement distributed inference topologies; Define technical direction and benchmarking strategies
Seniority
Senior, hands-on IC