Senior Software Engineer (vLLM)
Core
Deploy, extend, and optimize vLLM and other inference serving engines to support open-weight LLMs in regulated and high-assurance environments.
Role type
Senior IC software engineer (LLM inference infrastructure)
Builds
High-performance inference serving systems for open-weight models
Domain
AI infrastructure / LLM inference / Open source
Deliverable
production ML models
Required skills
Python, C/C++, Rust, CUDA, LLM inference mechanics (KV caching, continuous batching, quantisation), hardware accelerator optimization, open-source contribution, performance profiling
Preferred skills
Experience in highly regulated environments, AI safety research, MLOps practices
Technologies
vLLM, PyTorch, Hugging Face TGI, TensorRT-LLM, Ray
Responsibilities
Deploy and monitor open weight models served using vLLM; Implement new features within vLLM for novel hardware architectures; Collaborate with the open-source vLLM community to upstream core changes; Troubleshoot and optimize inference performance for latency, throughput, and hardware utilisation
Seniority
Senior, hands-on IC