硬件加速推理引擎运行时开发工程师-AI工具链
Core
Design and implement core runtime components for AI inference engines, including model loading, graph optimization, operator scheduling, and memory management for deep learning AI chips.
Role type
Senior IC hardware-accelerated inference engine runtime developer
Builds
Runtime/UMD software stacks for deep learning AI chips
Domain
AI hardware acceleration / Deep learning infrastructure
Deliverable
production ML models
Required skills
C++, computer architecture, data structures and algorithms, heterogeneous runtime development, driver development, CUDA Runtime, AMD ROCm/CLR
Preferred skills
GPU/NPU architecture, IC implementation details, AI fundamentals, Python
Responsibilities
Design and implement model loading, graph optimization, operator scheduling, and memory management components; Develop and maintain Runtime/UMD software stacks for AI chips.