LLM Inference Deployment Engineer
Core
Optimizing, deploying, and scaling large language models (LLMs) for high-performance inference on energy-efficient AI accelerators.
Role type
LLM Inference Deployment Engineer
Builds
High-performance inference pipelines and runtime execution environments for generative AI applications.
Domain
AI hardware and software systems for edge-to-cloud computing
Deliverable
production ML models
Required skills
LLM inference deployment, model optimization, runtime engineering, LLM inference frameworks (PyTorch, ONNX Runtime, vLLM, TensorRT-LLM, DeepSpeed), Python programming, containerized AI deployments (Docker, Kubernetes, Triton Inference Server, TensorFlow Serving, TorchServe), LLM memory optimization strategies, real-time LLM application development
Preferred skills
Experience with GPT, LLaMA, Mistral, Falcon models; experience with post-training libraries like HuggingFace; knowledge of batching, caching, and tensor parallelism
Technologies
ONNX Runtime, vLLM, Docker, Kubernetes, Triton Inference Server, TensorFlow Serving, TorchServe, HuggingFace
Responsibilities
Deploy and optimize LLMs post-training from libraries like HuggingFace; Utilize inference runtimes such as ONNX Runtime, vLLM for efficient execution; Optimize batching, caching, and tensor parallelism to improve LLM scalability in real-time applications; Develop and maintain high-performance inference pipelines using Docker, Kubernetes, and other inference servers
Seniority
Mid-to-Senior level IC