CareerPlanGet AI match score →

LLM Inference Deployment Engineer

Canada💼 Full-time💰 $180,000–$180,000🗓 2026-07-09 → 2026-07-31

Core

Optimizing, deploying, and scaling large language models (LLMs) for high-performance inference on energy-efficient AI accelerators.

Role type

LLM Inference Deployment Engineer

Builds

High-performance inference pipelines and runtime execution environments for generative AI applications.

Domain

AI hardware and software systems for edge-to-cloud computing

Deliverable

production ML models

Required skills

LLM inference deployment, model optimization, runtime engineering, LLM inference frameworks (PyTorch, ONNX Runtime, vLLM, TensorRT-LLM, DeepSpeed), Python programming, containerized AI deployments (Docker, Kubernetes, Triton Inference Server, TensorFlow Serving, TorchServe), LLM memory optimization strategies, real-time LLM application development

Preferred skills

Experience with GPT, LLaMA, Mistral, Falcon models; experience with post-training libraries like HuggingFace; knowledge of batching, caching, and tensor parallelism

Technologies

ONNX Runtime, vLLM, Docker, Kubernetes, Triton Inference Server, TensorFlow Serving, TorchServe, HuggingFace

Responsibilities

Deploy and optimize LLMs post-training from libraries like HuggingFace; Utilize inference runtimes such as ONNX Runtime, vLLM for efficient execution; Optimize batching, caching, and tensor parallelism to improve LLM scalability in real-time applications; Develop and maintain high-performance inference pipelines using Docker, Kubernetes, and other inference servers

Seniority

Mid-to-Senior level IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗