Research Engineer, Model Inference & Serving - London
Core
Build and operate the inference stack serving multimodal agentic models, optimizing latency, throughput, and cost while co-designing training-time decisions with the Models team.
Role type
Senior IC research engineer (model inference & serving)
Builds
Production inference systems for agentic AI products
Domain
AI / Machine Learning / Systems
Deliverable
production ML models
Required skills
Python, Rust/C++/Go, PyTorch/JAX, distributed systems, cloud environments, Kubernetes, transformers, multimodal architectures
Preferred skills
vLLM/SGLang/TensorRT-LLM, GPU kernel development (CUDA/Triton), edge inference, quantization, speculative decoding, KV-cache compression, multimodal models, agentic systems
Technologies
vLLM, SGLang, TensorRT-LLM, CUDA, Triton, llama.cpp, MLX, ONNX Runtime, Kubernetes, PyTorch, JAX
Responsibilities
Build and operate the inference stack for multimodal agentic models; Improve latency, throughput, and cost of model serving; Research and implement inference techniques for agent workloads; Co-design with the Models team on training-time decisions; Collaborate with cross-functional teams to integrate inference into products; Evaluate inference, serving, and hardware platforms; Stay current with advancements in inference and accelerator technology
Seniority
Senior, hands-on IC