Inference Optimization Engineer (local / edge runtime)
Core
Optimizing inference engines (llama.cpp, vLLM) for latency, throughput, and memory on constrained local and edge hardware (PCs, iGPUs) to enable private, low-cost AI agents.
Role type
Inference Optimization Engineer (local/edge runtime)
Builds
Efficient local AI inference engines for consumer PCs and edge devices
Domain
AI/ML inference, embedded systems, hardware optimization
Deliverable
production ML models
Required skills
C++, Python, systems-level code profiling, LLM inference internals (attention, KV cache, decoding), Linux build systems, low-level debugging
Preferred skills
llama.cpp, vLLM, ggml, GPU/accelerator programming (Vulkan, CUDA, SYCL, Metal), SIMD/CPU kernels, quantization formats (GGUF, AWQ, GPTQ), open-source contributions
Technologies
llama.cpp, vLLM, Vulkan, SYCL, oneAPI, CUDA, GGUF, AWQ, GPTQ
Responsibilities
Profile and optimize local inference for latency and memory; Tune KV cache and continuous batching for interactive workloads; Drive quantization strategy and validate quality; Cut CPU overhead and improve engine lifecycle; Benchmark across hardware tiers; Upstream fixes to open-source engines
Seniority
Senior, hands-on IC