CareerPlanSign in

硬件加速推理引擎运行时开发工程师-Data(西安)

西安💼 Full-time🗓 2026-09-28

Core

Design and implement core runtime components for an inference engine, including model loading, graph optimization, operator scheduling, and memory management.

Role type

Senior IC inference engine runtime engineer

Builds

High-performance inference runtime libraries supporting multiple deep learning frameworks

Domain

Deep Learning / High-Performance Computing / Hardware Acceleration

Deliverable

production ML models

Required skills

C++, Python, deep learning framework runtime mechanisms, computer architecture (CPU/GPU/NPU), multi-threading, memory management, performance optimization

Preferred skills

model quantization/pruning/distillation, LLVM/MLIR compiler technology, open source contributions, edge computing/embedded systems, inference engine development (TensorRT/OpenVINO/TVM)

Technologies

TensorFlow, PyTorch, ONNX, Linux, Windows, embedded systems, LLVM, MLIR

Responsibilities

Analyze and resolve runtime performance bottlenecks to improve throughput and reduce latency; develop cross-platform support for various operating systems and hardware architectures; develop and maintain the compilation toolchain for model conversion, quantization, and pruning; provide debugging and profiling tools; collaborate with algorithm and product teams for rapid integration and deployment of new models and operators.

Sourced via bytedance · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.