CareerPlanGet AI match score →

Senior Software Engineer (vLLM)

Cambridge, England, United Kingdom💼 Full-time🗓 2026-06-26 → 2026-07-31

Core

Deploy, extend, and optimize vLLM and other inference serving engines to support open-weight LLMs in regulated and high-assurance environments.

Role type

Senior IC software engineer (LLM inference infrastructure)

Builds

High-performance inference serving systems for open-weight models

Domain

AI infrastructure / LLM inference / Open source

Deliverable

production ML models

Required skills

Python, C/C++, Rust, CUDA, LLM inference mechanics (KV caching, continuous batching, quantisation), hardware accelerator optimization, open-source contribution, performance profiling

Preferred skills

Experience in highly regulated environments, AI safety research, MLOps practices

Technologies

vLLM, PyTorch, Hugging Face TGI, TensorRT-LLM, Ray

Responsibilities

Deploy and monitor open weight models served using vLLM; Implement new features within vLLM for novel hardware architectures; Collaborate with the open-source vLLM community to upstream core changes; Troubleshoot and optimize inference performance for latency, throughput, and hardware utilisation

Seniority

Senior, hands-on IC

Sourced via workable · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workable ↗