深度学习模型量化与压缩_XC
Core
Design, implement, and optimize deep learning models for deployment on embedded systems, specifically focusing on quantization, compression, and performance tuning for autonomous driving applications.
Role type
Senior Deep Learning Model Optimization Engineer (Embedded Systems)
Builds
Optimized deep learning models for embedded chips (NVIDIA, Qualcomm, Horizon)
Domain
Autonomous Driving / Embedded AI
Deliverable
production ML models
Required skills
Deep learning model quantization, Post-training Quantization, Quantization-Aware Training, Pruning, Knowledge Distillation, Low-rank decomposition, Embedded system deployment, Hardware accelerator optimization (GPU/NPU), Python, C++, PyTorch, TensorFlow
Preferred skills
NVIDIA TensorRT, TVM, PyTorch FX, Real-time systems, ADAS systems, Perception algorithms, Horizon chips, Qualcomm chips
Technologies
Python, C++, PyTorch, TensorFlow, NVIDIA TensorRT, TVM, PyTorch FX
Responsibilities
Develop and implement cutting-edge quantization and compression techniques; Balance model accuracy, size, and inference speed based on experiments; Deploy models to embedded platforms and optimize for specific hardware accelerators; Analyze model performance to identify bottlenecks and improve real-time inference capabilities.
Seniority
Senior, hands-on IC