Python Inference Engineer
Core
Design and deliver the inference layer for the Gcore Inference platform, integrating frameworks to bring multimodal models into production.
Role type
Senior IC Python Inference Engineer
Builds
Production inference platforms and features for AI-driven digital experiences
Domain
Cloud infrastructure, GPU computing, and AI model deployment
Deliverable
production ML models
Required skills
Python, PyTorch, Linux, Docker, Kubernetes, distributed systems, GPU computing, model optimization, cluster scheduling, debugging complex software/hardware issues
Preferred skills
vLLM, SGLang, NVIDIA Dynamo, TensorRT-LLM, CUDA, Triton, TensorRT, quantization, continuous batching, speculative decoding, prefix caching, chunked prefill, LoRA serving, profiling model latency/throughput/memory/GPU utilization, distributed inference, multi-GPU systems, autoscaling, open-source contributions
Technologies
vLLM, SGLang, NVIDIA Dynamo, TensorRT-LLM, PyTorch, Kubernetes, Docker, CUDA, Triton, TensorRT
Responsibilities
Build and improve the inference layer; Integrate and operate inference frameworks; Bring new language and multimodal models into production; Improve inference latency, throughput, memory use, GPU utilization, and cost efficiency; Debug performance and reliability issues across model code, inference frameworks, GPU execution, networking, and Kubernetes; Contribute improvements to open-source inference projects
Seniority
Senior, hands-on IC