Senior Machine Learning Engineer, LLM Inference Optimization
Core
Drive optimization of large language and vision-language model inference from model artifacts through production deployment to improve latency, throughput, memory efficiency, GPU utilization, reliability, and cost per token.
Role type
Senior IC machine learning engineer (LLM inference optimization)
Builds
High-throughput AI inference serving systems for production environments
Domain
AI Infrastructure / Large Language Models / GPU Computing
Deliverable
production ML models
Required skills
Python, PyTorch, LLM/VLM inference optimization, modern inference engines (vLLM, SGLang, TensorRT-LLM, Triton), transformer inference bottlenecks analysis, quantitative performance trade-off reasoning, complex performance problem diagnosis
Preferred skills
Model compression workflows (quantization, distillation), advanced inference techniques (speculative decoding, KV-cache optimization), agentic workload support, CUDA/Triton, open-source contributions
Technologies
vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, PyTorch, CUDA
Responsibilities
Own optimization initiatives for specific model families and inference serving backends; Evaluate inference engines and recommend serving configurations; Diagnose and resolve model quality, performance, and reliability regressions; Optimize LLM/VLM endpoints for latency, throughput, and cost; Deploy and benchmark modern inference engines; Build and productionize model-compression workflows; Implement advanced inference techniques like speculative decoding and continuous batching; Develop reproducible benchmark harnesses; Partner with GPU kernel and platform engineers to identify bottlenecks; Produce design documentation and performance reports; Contribute to safe production rollouts
Seniority
Senior, hands-on IC