昆仑芯-深度学习高性能计算研发工程师(J90411)
Core
Develop high-performance deep learning computing libraries for Kunlunxin AI chips to support various AI scenarios.
Role type
Senior IC deep learning high-performance computing engineer
Builds
High-performance computing libraries, distributed training systems, and chip interconnect architectures
Domain
AI hardware, deep learning frameworks, high-performance computing
Deliverable
production ML models
Required skills
C/C++, Python, Linux development, CUDA, OpenCL, deep learning frameworks (PyTorch, TensorFlow, PaddlePaddle)
Preferred skills
Recommendation systems, speech, NLP, vision application optimization
Technologies
CUDA, OpenCL, PaddlePaddle, PyTorch, TensorFlow
Responsibilities
Develop high-performance deep learning computing libraries; explore new AI chip programming models and architectures; optimize PaddlePaddle graph performance; optimize large-scale distributed training and develop AI chip communication libraries
Seniority
Senior, hands-on IC