Inference Engineer
Core
Own the cost and performance of the inference stack to efficiently serve models under varying workloads and traffic.
Role type
Senior IC inference systems engineer (via careerplan.io/jobs/8dae2814-dcde-4dd4-ae08-720825b8ae7e-inference-engineer-at-adaption)
Builds
High-throughput, low-latency model serving infrastructure
Domain
AI/ML inference systems and performance engineering
Deliverable
production ML models
Required skills
KV-cache management, continuous batching, speculative decoding, quantization, model serving (prefill/decode), memory bandwidth optimization, concurrency, systems programming (Python/C++/Rust), GPU performance (CUDA/NCCL/kernels), profiling and measurement systems
Preferred skills
Routing optimization between infrastructure and external providers, kernel-level optimization
Technologies
vLLM, SGLang, TensorRT-LLM
Responsibilities
Improve throughput, cost, and tail latency via caching, batching, quantization, and decoding optimizations; Optimize long-context prefill and decode workloads based on production traffic; Tune routing between internal infrastructure and external providers; Build profiling systems to analyze time, memory, and compute usage; Work below the framework level in serving engines when necessary
Seniority
Senior, hands-on IC
