AI大模型优化工程师 - 音视频技术
Core
Optimizing inference performance of self-developed and open-source LLM/VLM multimodal models across heterogeneous chip architectures to support massive-scale computing for video, live streaming, and offline data processing.
Role type
Senior IC AI model inference optimization engineer (multimodal/heterogeneous chips)
Builds
High-performance inference services for video, live streaming, and large-scale offline data processing using tens of thousands of GPU/NPU cards.
Domain
AI/ML inference optimization, heterogeneous computing, video streaming infrastructure
Deliverable
production ML models
Required skills
C/C++, Python, Linux, computer architecture, parallel computing, GPU/NPU hardware architecture, CUDA, CUTLASS, AscendC, BangC, LLM/VLM model structures, inference frameworks (vLLM, SGLang), model parallelism strategies, operator optimization, quantization algorithms
Preferred skills
Research on cutting-edge inference acceleration techniques, hardware-software co-optimization, new heterogeneous hardware characteristics
Technologies
CUDA, CUTLASS, AscendC, BangC, vLLM, SGLang, heterogeneous chips (GPU/NPU)
Responsibilities
Optimize inference of self-developed and open-source LLM/VLM models on multiple heterogeneous chips; Improve the adaptation and optimization efficiency of the inference technology stack for different chips; Analyze and evaluate new heterogeneous chips for various large models; Research and implement cutting-edge technologies such as inference acceleration and hardware-software co-optimization.
Seniority
Senior, hands-on IC