CareerPlanGet AI match score →

MLOps Engineer (LLM/GenAI)

Sheffield, England, UK💼 Full-time🗓 2026-07-07 → 2026-08-01

Core

Design, build, and operate scalable model hosting platforms for LLMs, embeddings, and speech technologies, optimizing inference performance and managing end-to-end fine-tuning pipelines.

Role type

MLOps Engineer (LLM/GenAI)

Builds

Scalable model hosting platforms for LLMs, embeddings, and STT/TTS

Domain

Financial services (HSBC) / Large Language Models / Inference Optimization

Deliverable

production ML models

Required skills

Python, CUDA, GPU/CPU architecture, HPC fundamentals, inference optimization (KV-cache, batching, quantization), framework integration (vLLM, TensorRT-LLM, SGLang), Docker, Kubernetes, cloud platforms (AWS/GCP/Azure), distributed training, hyperparameter tuning, LoRA/QLoRA

Preferred skills

LLM experience

Technologies

vLLM, TensorRT-LLM, SGLang, Docker, Kubernetes, AWS, GCP, Azure, HF, Accelerate

Responsibilities

Design and operate scalable model hosting platforms for LLMs, embeddings, and STT/TTS; Optimize inference for latency, throughput, and cost; Evaluate and integrate inference frameworks; Own inference health/performance monitoring and troubleshoot bottlenecks; Build end-to-end fine-tuning pipelines and integrate fine-tuned models into the hosting/inference stack

Seniority

Mid-to-Senior, hands-on IC

Sourced via linkedin · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on LinkedIn ↗