AI Inference Engineer QVAC (100% remote Worldwide)
Core
Own the C++ inference backbone for QVAC's local AI stack, enabling fast, reliable, on-device AI execution without cloud dependency.
Role type
Senior IC C++ inference engineer (edge AI)
Builds
Production-ready inference engines for private, on-device AI experiences
Domain
Fintech + Edge AI / On-device ML
Deliverable
production ML models
Required skills
C++, llama.cpp, ggml, GPU programming (CUDA/Vulkan/Metal/OpenCL), deep learning concepts, transformer/LLM/diffusion model architectures
Preferred skills
LLM fine-tuning, model productionization, new model architecture research, distributed systems, JavaScript
Technologies
llama.cpp, ggml, CUDA, Vulkan, Metal, OpenCL
Responsibilities
Deploy ML models to edge devices using llama.cpp and ggml; Collaborate with researchers to transition models from research to production; Integrate AI features into existing products
Seniority
Senior, hands-on IC