CareerPlanSign in

Inference Engineer

San Francisco💼 Full-time🗓 2026-10-05 → 2026-10-07

Core

Own the cost and performance of the inference stack to efficiently serve models under varying workloads and traffic.

Builds

High-throughput, low-latency model serving infrastructure

Domain

AI/ML inference systems and performance engineering

Deliverable

production ML models

Required skills

KV-cache management, continuous batching, speculative decoding, quantization, model serving (prefill/decode), memory bandwidth optimization, concurrency, systems programming (Python/C++/Rust), GPU performance (CUDA/NCCL/kernels), profiling and measurement systems

Preferred skills

Routing optimization between infrastructure and external providers, kernel-level optimization

Technologies

vLLM, SGLang, TensorRT-LLM

Responsibilities

Improve throughput, cost, and tail latency via caching, batching, quantization, and decoding optimizations; Optimize long-context prefill and decode workloads based on production traffic; Tune routing between internal infrastructure and external providers; Build profiling systems to analyze time, memory, and compute usage; Work below the framework level in serving engines when necessary

Seniority

Senior, hands-on IC