LLM/VLM大模型推理算子/框架优化专家 - Data AML
Core
Optimizing inference for LLMs and VLMs on heterogeneous chips to maximize large-scale compute power for the Volcano Engine Ark platform.
Role type
Senior IC LLM/VLM inference optimization engineer
Builds
High-performance inference systems for Deepseek, Kimi, GLM, and other large models on diverse chip architectures
Domain
AI Infrastructure / Heterogeneous Computing / Large Language Models
Deliverable
production ML models
Required skills
C/C++, Python, Linux, Computer Architecture, Parallel Computing, GPU/NPU Hardware Architecture, Inference Optimization Software Stacks (CUDA, CUTLASS, AscendC, BangC, HIP, FlyDSL), Heterogeneous Chip AI Model Performance Analysis, Operator Optimization
Preferred skills
LLM/VLM Architecture Knowledge, Inference Frameworks (vLLM, SGLang), Large Model Parallel Strategies, Quantization Algorithms
Technologies
CUDA, CUTLASS, AscendC, BangC, HIP, FlyDSL, vLLM, SGLang
Responsibilities
Optimize inference for LLMs/VLMs across multiple heterogeneous chips; Enhance the inference adaptation and optimization technology system for different chips; Analyze and evaluate new heterogeneous chips for large models; Research and implement cutting-edge inference acceleration and hardware-software co-optimization techniques.