CareerPlanSign in

Python Inference Engineer

🌐 Remote💼 Full-time🗓 2026-09-10 → 2026-09-25

Core

Design and deliver the inference layer for the Gcore Inference platform, integrating frameworks to bring multimodal models into production.

Role type

Senior IC Python Inference Engineer

Builds

Production inference platforms and features for AI-driven digital experiences

Domain

Cloud infrastructure, GPU computing, and AI model deployment

Deliverable

production ML models

Required skills

Python, PyTorch, Linux, Docker, Kubernetes, distributed systems, GPU computing, model optimization, cluster scheduling, debugging complex software/hardware issues

Preferred skills

vLLM, SGLang, NVIDIA Dynamo, TensorRT-LLM, CUDA, Triton, TensorRT, quantization, continuous batching, speculative decoding, prefix caching, chunked prefill, LoRA serving, profiling model latency/throughput/memory/GPU utilization, distributed inference, multi-GPU systems, autoscaling, open-source contributions

Technologies

vLLM, SGLang, NVIDIA Dynamo, TensorRT-LLM, PyTorch, Kubernetes, Docker, CUDA, Triton, TensorRT

Responsibilities

Build and improve the inference layer; Integrate and operate inference frameworks; Bring new language and multimodal models into production; Improve inference latency, throughput, memory use, GPU utilization, and cost efficiency; Debug performance and reliability issues across model code, inference frameworks, GPU execution, networking, and Kubernetes; Contribute improvements to open-source inference projects

Seniority

Senior, hands-on IC

Sourced via codingjobboard · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.