Edge AI Model Optimization Software Engineer
Core
Design and implement production-grade optimization tools and workflows to enable efficient execution of Generative AI, Transformers, and VLMs on resource-constrained edge platforms.
Role type
Senior IC AI Optimization Engineer (Edge AI)
Builds
High-performance software toolchains for on-device GenAI execution
Domain
Edge AI, Embedded Systems, Deep Learning
Deliverable
production ML models
Required skills
Neural Network Quantization, Mixed-precision flows, PTQ/QAT workflows, Python, C/C++, PyTorch, ONNX, CNN architectures, Generative AI (Transformers), Numerical approximation algorithms, Hardware-aware optimization
Preferred skills
Hardware accelerators, Device-level profiling, Embedded systems constraints, State-of-the-art quantization (GPTQ, Smoothquant), MLIR, TVM
Technologies
PyTorch, ONNX, Ara 2, PTQ, QAT, MLIR, TVM
Responsibilities
Design and implement quantization features including mixed-precision flows; Maintain scalable PTQ and QAT workflows; Evaluate novel quantization techniques and engineer production deployment recipes; Implement approximation algorithms ensuring bit-exactness on target hardware; Profile and optimize hot paths of the optimization toolchain; Act as technical bridge between AI Research and Hardware Engineering; Document algorithmic tradeoffs and derive gold-standard deployment recipes
Seniority
Senior, hands-on IC