CareerPlanGet AI match score →

Inference Optimization Engineer (local / edge runtime)

4 Locations💼 Full-time💰 $170,500–$315,490🗓 2026-06-15 → 2026-07-30

Core

Optimizing inference engines (llama.cpp, vLLM) for latency, throughput, and memory on constrained local and edge hardware (PCs, iGPUs) to enable private, low-cost AI agents.

Role type

Inference Optimization Engineer (local/edge runtime)

Builds

Efficient local AI inference engines for consumer PCs and edge devices

Domain

AI/ML inference, embedded systems, hardware optimization

Deliverable

production ML models

Required skills

C++, Python, systems-level code profiling, LLM inference internals (attention, KV cache, decoding), Linux build systems, low-level debugging

Preferred skills

llama.cpp, vLLM, ggml, GPU/accelerator programming (Vulkan, CUDA, SYCL, Metal), SIMD/CPU kernels, quantization formats (GGUF, AWQ, GPTQ), open-source contributions

Technologies

llama.cpp, vLLM, Vulkan, SYCL, oneAPI, CUDA, GGUF, AWQ, GPTQ

Responsibilities

Profile and optimize local inference for latency and memory; Tune KV cache and continuous batching for interactive workloads; Drive quantization strategy and validate quality; Cut CPU overhead and improve engine lifecycle; Benchmark across hardware tiers; Upstream fixes to open-source engines

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗