CareerPlanSign in

Inference Performance Engineer

San Francisco💼 Full-time🗓 2026-08-21 → 2026-09-26

Core

Own the cost and performance of the inference stack to efficiently serve models under varying workloads and traffic.

Role type

Senior IC inference performance engineer

Builds

High-throughput, low-latency model serving systems

Domain

AI/ML inference infrastructure

Deliverable

production ML models

Required skills

KV-cache management, continuous batching, speculative decoding, quantization, long-context optimization, routing strategy, profiling systems, vLLM, SGLang, TensorRT-LLM, Python, C++, Rust, GPU performance, CUDA, NCCL, mixed precision, kernel optimization

Preferred skills

None stated

Technologies

vLLM, SGLang, TensorRT-LLM, CUDA, NCCL

Responsibilities

Improve throughput, cost, and tail latency via caching, batching, and quantization; Optimize prefill and decode workloads based on production traffic; Tune routing between internal infrastructure and external providers; Build profiling and measurement systems for time, memory, and compute analysis

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.