CareerPlanSign in

Research Engineer - Inference

United Kingdom🌐 Remote💼 Full-time🗓 2026-08-28 → 2026-09-26

Core

Deploy and optimize frontier AI models in production for real-time, streaming workloads, turning research breakthroughs into scalable products.

Role type

Senior IC research engineer (inference optimization)

Builds

High-performance serving systems and tooling for real-time AI audio models

Domain

AI/ML inference, real-time streaming systems

Deliverable

production ML models

Required skills

GPU programming, inference optimization, CUDA, Triton, TensorRT, vLLM, SGLang, profiling, bottleneck elimination, custom kernel development

Preferred skills

Experience with latency-sensitive applications, model quantization, distillation, KV-cache optimization, batching strategies

Technologies

CUDA, Triton, TensorRT, vLLM, SGLang

Responsibilities

Deploy state-of-the-art models to production; optimize inference performance across the stack; build and tune high-performance serving systems; create tooling for researchers to ship models quickly

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.