CareerPlanSign in

硬件加速推理引擎运行时开发工程师-AI工具链

上海💼 Full-time🗓 2026-09-28

Core

Design and implement core runtime components for AI inference engines, including model loading, graph optimization, operator scheduling, and memory management for deep learning AI chips.

Role type

Senior IC hardware-accelerated inference engine runtime developer

Builds

Runtime/UMD software stacks for deep learning AI chips

Domain

AI hardware acceleration / Deep learning infrastructure

Deliverable

production ML models

Required skills

C++, computer architecture, data structures and algorithms, heterogeneous runtime development, driver development, CUDA Runtime, AMD ROCm/CLR

Preferred skills

GPU/NPU architecture, IC implementation details, AI fundamentals, Python

Responsibilities

Design and implement model loading, graph optimization, operator scheduling, and memory management components; Develop and maintain Runtime/UMD software stacks for AI chips.

Sourced via bytedance · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.