Senior ML Engineer - Kimchi (LLM Inference Optimization)
Core
Optimizing LLM inference throughput, latency, and KV cache utilization on cloud infrastructure to reduce costs and improve performance for customers.
Role type
Senior IC machine-learning engineer (LLM inference optimization)
Builds
Autonomous decision-making layer for matching workloads to cost-efficient LLM configurations and serving settings
Domain
Cloud-native AI infrastructure, Kubernetes, Large Language Model serving
Deliverable
production ML models
Required skills
Python (production services), vLLM/SGLang/TensorRT-LLM, quantization tradeoffs, distributed systems (collective communication, sharding), measurement-driven optimization, kernel-level tuning
Preferred skills
None stated
Technologies
vLLM, SGLang, TensorRT-LLM, PyTorch, CUDA, Kubernetes, gRPC, ClickHouse, PostgreSQL, GCP Pub/Sub, AWS, Azure, GitLab CI, ArgoCD, Prometheus, Grafana, Loki, Tempo
Responsibilities
Tune kernels and schedulers to maximize GPU throughput; profile and fix latency bottlenecks (TTFT, TPOT); implement KV cache optimization strategies (paged attention, prefix caching); perform empirical quantization without quality regression; optimize memory footprint and cold starts; design distributed inference topologies; set technical direction and benchmarking standards
Seniority
Senior, hands-on IC