Principal Edge AI Software Architect
Core
Design and implement advanced machine learning solutions, specifically optimizing and deploying Large Language Models (LLMs) on resource-constrained edge devices and embedded systems.
Role type
Principal Edge AI Software Architect
Builds
Scalable Edge AI inference engines, optimized model compression pipelines, and custom kernels for edge AI accelerators.
Domain
Embedded systems, Edge AI, Large Language Models (LLMs)
Deliverable
production ML models
Required skills
LLM optimization (quantization, pruning, distillation, LoRA, QLoRA), C/C++ and Python programming, embedded software development, real-time systems, computer architecture, memory hierarchies, hardware acceleration
Preferred skills
TensorFlow, PyTorch, ONNX, TensorFlow Lite, Pytorch Mobile, TensorRT, OpenVINO, NPU/DSP/GPU profiling, multimodal models (vision-language, audio-text), security considerations for edge AI
Technologies
TensorFlow, PyTorch, ONNX, TensorFlow Lite, PyTorch Mobile, TensorRT, OpenVINO, C, C++, Python
Responsibilities
Design scalable Edge AI inference engines for microcontrollers and embedded systems; Define technical roadmaps for deploying LLMs on edge hardware; Lead architecture of model compression and optimization pipelines; Optimize and deploy LLMs using quantization, pruning, and knowledge distillation; Develop custom kernels for edge AI accelerators; Train, fine-tune, and optimize ML models for edge deployment; Profile and optimize model performance on edge AI accelerators to meet latency and power constraints
Seniority
Principal, hands-on IC with strategic roadmap definition