硬件加速推理引擎运行时开发工程师-Data(西安)
Core
Design and implement core runtime components for an inference engine, including model loading, graph optimization, operator scheduling, and memory management.
Role type
Senior IC inference engine runtime engineer
Builds
High-performance inference runtime libraries supporting multiple deep learning frameworks
Domain
Deep Learning / High-Performance Computing / Hardware Acceleration
Deliverable
production ML models
Required skills
C++, Python, deep learning framework runtime mechanisms, computer architecture (CPU/GPU/NPU), multi-threading, memory management, performance optimization
Preferred skills
model quantization/pruning/distillation, LLVM/MLIR compiler technology, open source contributions, edge computing/embedded systems, inference engine development (TensorRT/OpenVINO/TVM)
Technologies
TensorFlow, PyTorch, ONNX, Linux, Windows, embedded systems, LLVM, MLIR
Responsibilities
Analyze and resolve runtime performance bottlenecks to improve throughput and reduce latency; develop cross-platform support for various operating systems and hardware architectures; develop and maintain the compilation toolchain for model conversion, quantization, and pruning; provide debugging and profiling tools; collaborate with algorithm and product teams for rapid integration and deployment of new models and operators.