大模型异构芯片推理适配调优工程师(深圳/北京/上海/杭州)
Core
Lead the development and implementation of distributed heterogeneous inference systems for large language models across GPGPU/NPU/XPU chips, focusing on low latency and high throughput.
Role type
Senior IC machine-learning systems engineer (heterogeneous inference)
Builds
Distributed inference systems for large language models (e.g., DeepSeek, Qwen)
Domain
AI Infrastructure / Large Language Models / Heterogeneous Computing
Deliverable
production ML models
Required skills
C/C++, Python, CUDA, Triton, AscendC, BangC, NCCL, NVLink, RoCE, GPGPU architecture, NPU architecture, XPU architecture, distributed system development, model adaptation, precision tuning, system performance profiling, operator optimization, inference engine optimization (vllm/sglang), PD separation architecture
Preferred skills
Experience with multiple chip architectures (AMD, Hygon, Moore Threads,沐曦, Ascend, Cambricon, Kunlun), deep experience with LLM structures (DeepSeek, Qwen), large EP compute fusion operator optimization, key operator optimization (Attention MLA/GQA, Sparse Attention DSA, Linear Attention, GEMM, Group GEMM)
Responsibilities
Adapt mainstream and self-developed large models to multiple heterogeneous chips; resolve precision anomalies in adaptation; profile and optimize the full inference chain for performance; optimize inference framework/engine architecture for multi-chip characteristics; develop and tune core operators for specific chip micro-architectures.
Seniority
Senior, hands-on IC