异构加速框架工程师(深圳/北京/上海)
Core
Design and optimize GPU/AI chip performance for AI inference, collaborating with algorithm teams to build high-performance operators and framework layers.
Role type
Senior IC machine-learning framework engineer (GPU/AI chip optimization)
Builds
High-performance inference frameworks and optimized AI operators
Domain
AI Infrastructure / High-Performance Computing
Deliverable
production ML models
Required skills
C/C++, Python, CUDA, Triton, Ascend C, Cublas, Cutlass, CK, Torch-Compile, parallel computing, memory optimization, communication optimization
Preferred skills
Dynamic computation graph compilation, MOE models, KV Cache optimization, dynamic batching, custom Attention operators
Technologies
CUDA, Triton, Ascend C, Cublas, Cutlass, CK, Torch-Compile
Responsibilities
Co-design GPU/AI chip performance optimizations with algorithm teams; Innovate and optimize core modules in ML frameworks; Design and implement high-performance operators and feature enablement components; Explore frontier technologies like MOE and dynamic graph compilation.