CareerPlanSign in

Inference Performance & Deployment - Member of Technical Staff

London💼 Full-time🗓 2026-07-22 → 2026-09-26

Core

Design tooling and methodologies to ground heterogeneous AI infrastructure in real-world performance, acting as the integration point between engineering functions and production environments.

Role type

Member of Technical Staff (Inference Performance & Deployment)

Builds

Deployment patterns, benchmarking harnesses, regression suites, and performance dashboards for heterogeneous compute systems.

Domain

AI Infrastructure / Heterogeneous Compute / Inference Systems

Deliverable

production ML models

Required skills

Large model inference deployment, multi-node GPU deployment, end-to-end performance characterization, serving frameworks (Dynamo, Triton), system bottleneck isolation, reproducible measurement methodology

Preferred skills

Cloud instance self-hosting, orchestration and routing software, caching and request scheduling, resource allocation

Technologies

Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, Mixx, Dynamo, Triton Inference Server

Responsibilities

Run experiments self-hosting models across providers and hardware configurations, develop optimized deployment patterns, work on orchestration and routing software, act as integration point for accelerator support and engine features, build and maintain benchmarking harnesses and dashboards

Seniority

Mid-Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.