Senior Machine Learning Engineer, LLM Inference Optimization
Core
Building and optimizing large-scale LLM/VLM inference infrastructure for training and post-training capabilities, focusing on latency, throughput, and cost efficiency.
Role type
Senior IC Machine Learning Engineer (LLM Inference Optimization)
Builds
Production-grade inference engines, model compression workflows, and benchmark harnesses for frontier AI models.
Domain
Cloud Infrastructure / Large Language Models / Distributed Systems
Deliverable
production ML models
Required skills
Python, PyTorch, LLM/VLM inference optimization, transformer inference bottlenecks, quantitative performance analysis, inference engine deployment, model compression techniques
Preferred skills
Quantization-aware training, speculative decoding, agentic workloads, CUDA/Triton, open-source contributions to inference stacks
Technologies
vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, FlashInfer, LMCache
Responsibilities
Own optimization work for specific model families and serving backends; Run engine comparisons and recommend serving configurations; Debug model quality or performance regressions; Optimize LLM/VLM endpoints for latency, throughput, and cost; Deploy and extend inference engines; Build model-compression workflows; Implement speculative decoding and KV-cache optimizations; Build reproducible benchmark harnesses; Partner with GPU kernel and platform engineers.
Seniority
Senior, hands-on IC