火山引擎机器学习异构硬件开发工程师-Data AML
Core
Develop and optimize high-performance operators and compilation stacks for heterogeneous computing chips to accelerate machine learning inference and training workloads.
Role type
Senior IC heterogeneous computing engineer (ML hardware)
Builds
Optimized inference and training pipelines for ML models on custom chips
Domain
AI Infrastructure / Heterogeneous Computing
Deliverable
production ML models
Required skills
C/C++, Python, PyTorch, deep learning model architecture, parallel computing architectures, high-performance operator development, chip-specific optimization, memory management, compiler technology
Preferred skills
Experience with domestic chips (Cambricon, Ascend), GPU architecture (CUDA, cuBLAS), AI Compiler stacks (MLIR, Torch2.0+, Triton), SIMD/SIMT models, large-scale training, compute-fusion
Responsibilities
Evaluate heterogeneous computing chips and build assessment frameworks; adapt chips for inference to reduce latency and increase throughput; optimize memory usage and throughput for training; develop high-performance operators; implement efficient heterogeneous hardware programming paradigms via compilation.