CareerPlanSign in

Senior Machine Learning Engineer, LLM Inference Optimization

💼 Full-time🗓 2026-09-23 → 2026-09-26

Core

Building fast, reliable, and cost-efficient inference services for frontier LLM/VLM models, optimizing from model artifacts through production deployment.

Role type

Senior Machine Learning Engineer (LLM Inference Optimization)

Builds

Production inference engines and serving backends for frontier AI models

Domain

Cloud Infrastructure / Large Language Models / Inference Optimization

Deliverable

production ML models

Required skills

Python, PyTorch, LLM/VLM inference optimization, transformer inference bottlenecks, latency/throughput/cost tradeoff analysis, inference engine configuration, model compression workflows, benchmarking, system design

Preferred skills

Quantization-aware training, post-training quantization, speculative decoding, agentic workloads, CUDA/Triton, open-source contributions

Technologies

vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, FlashInfer, LMCache

Responsibilities

Own optimization work for specific model families or serving backends; Run engine comparisons and recommend serving configurations; Debug model quality or performance regressions; Optimize endpoints for latency, throughput, memory efficiency, and cost; Deploy and extend inference engines; Build model-compression workflows; Implement speculative decoding and KV-cache optimizations; Build reproducible benchmark harnesses; Partner with kernel and platform engineers to diagnose bottlenecks

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.