Senior Machine Learning Engineer, LLM Inference Optimization
Core
Building fast, reliable, and cost-efficient inference services for frontier LLM/VLM models, optimizing from model artifacts through production deployment.
Role type
Senior Machine Learning Engineer (LLM Inference Optimization)
Builds
Production inference engines and serving backends for frontier AI models
Domain
Cloud Infrastructure / Large Language Models / Inference Optimization
Deliverable
production ML models
Required skills
Python, PyTorch, LLM/VLM inference optimization, transformer inference bottlenecks, latency/throughput/cost tradeoff analysis, inference engine configuration, model compression workflows, benchmarking, system design
Preferred skills
Quantization-aware training, post-training quantization, speculative decoding, agentic workloads, CUDA/Triton, open-source contributions
Technologies
vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, FlashInfer, LMCache
Responsibilities
Own optimization work for specific model families or serving backends; Run engine comparisons and recommend serving configurations; Debug model quality or performance regressions; Optimize endpoints for latency, throughput, memory efficiency, and cost; Deploy and extend inference engines; Build model-compression workflows; Implement speculative decoding and KV-cache optimizations; Build reproducible benchmark harnesses; Partner with kernel and platform engineers to diagnose bottlenecks
Seniority
Senior, hands-on IC