Inference Performance Engineer
Core
Own the cost and performance of the inference stack to efficiently serve models under varying workloads and traffic.
Role type
Senior IC inference performance engineer
Builds
High-throughput, low-latency model serving systems
Domain
AI/ML inference infrastructure
Deliverable
production ML models
Required skills
KV-cache management, continuous batching, speculative decoding, quantization, long-context optimization, routing strategy, profiling systems, vLLM, SGLang, TensorRT-LLM, Python, C++, Rust, GPU performance, CUDA, NCCL, mixed precision, kernel optimization
Preferred skills
None stated
Technologies
vLLM, SGLang, TensorRT-LLM, CUDA, NCCL
Responsibilities
Improve throughput, cost, and tail latency via caching, batching, and quantization; Optimize prefill and decode workloads based on production traffic; Tune routing between internal infrastructure and external providers; Build profiling and measurement systems for time, memory, and compute analysis
Seniority
Senior, hands-on IC