LLM Serving Engineer (Cloud AI Engineering), Senior / Staff Engineer
Core
Building scalable LLM inference platforms and optimizing serving packages for commercial deployment.
Role type
Senior/Staff LLM Serving Engineer (Cloud AI Engineering)
Builds
Scalable LLM inference platforms and serving packages (vLLM, SGLang, Triton-Inference server, etc.)
Domain
Cloud AI, Large Language Models, Inference Acceleration
Deliverable
production ML models
Required skills
LLM serving/orchestration packages, transformer-based architectures, PyTorch, distributed systems, computer architecture, ML accelerators, Python, deep learning workload optimization
Preferred skills
Open-source GenAI contributions, large-scale distributed systems architecture, high-level kernel design (PyTorch, CUDA, Triton), torch.compile/torchDynamo
Technologies
vLLM, SGLang, Triton-Inference server, Ollama, llm-d, KServe, LMCache, MoonCake, PyTorch, CUDA, Triton
Responsibilities
Building scalable LLM inference platforms using techniques like disaggregated serving, KV-Cache management, and speculative algorithms; Contributing to the development of LLM Serving packages; Collaborating with internal compiler, firmware, and platform teams to drive customer solutions; Identifying optimization opportunities in advanced algorithms like attention mechanisms and MoEs; Driving efficient serving through autoscaling, load balancing, and routing; Engaging with open-source serving communities to evolve frameworks
Seniority
Senior/Staff, hands-on IC with strategic thinking