MLOps Engineer (LLM/GenAI)
Core
Design, build, and operate scalable model hosting platforms for LLMs, embeddings, and speech technologies, optimizing inference performance and managing end-to-end fine-tuning pipelines.
Role type
MLOps Engineer (LLM/GenAI)
Builds
Scalable model hosting platforms for LLMs, embeddings, and STT/TTS
Domain
Financial services (HSBC) / Large Language Models / Inference Optimization
Deliverable
production ML models
Required skills
Python, CUDA, GPU/CPU architecture, HPC fundamentals, inference optimization (KV-cache, batching, quantization), framework integration (vLLM, TensorRT-LLM, SGLang), Docker, Kubernetes, cloud platforms (AWS/GCP/Azure), distributed training, hyperparameter tuning, LoRA/QLoRA
Preferred skills
LLM experience
Technologies
vLLM, TensorRT-LLM, SGLang, Docker, Kubernetes, AWS, GCP, Azure, HF, Accelerate
Responsibilities
Design and operate scalable model hosting platforms for LLMs, embeddings, and STT/TTS; Optimize inference for latency, throughput, and cost; Evaluate and integrate inference frameworks; Own inference health/performance monitoring and troubleshoot bottlenecks; Build end-to-end fine-tuning pipelines and integrate fine-tuned models into the hosting/inference stack
Seniority
Mid-to-Senior, hands-on IC