CareerPlanSign in

大模型异构芯片推理适配调优工程师(深圳/北京/上海/杭州)

Beijing, China💼 Full-time🗓 2026-09-28

Core

Lead the development and implementation of distributed heterogeneous inference systems for large language models across GPGPU/NPU/XPU chips, focusing on low latency and high throughput.

Role type

Senior IC machine-learning systems engineer (heterogeneous inference)

Builds

Distributed inference systems for large language models (e.g., DeepSeek, Qwen)

Domain

AI Infrastructure / Large Language Models / Heterogeneous Computing

Deliverable

production ML models

Required skills

C/C++, Python, CUDA, Triton, AscendC, BangC, NCCL, NVLink, RoCE, GPGPU architecture, NPU architecture, XPU architecture, distributed system development, model adaptation, precision tuning, system performance profiling, operator optimization, inference engine optimization (vllm/sglang), PD separation architecture

Preferred skills

Experience with multiple chip architectures (AMD, Hygon, Moore Threads,沐曦, Ascend, Cambricon, Kunlun), deep experience with LLM structures (DeepSeek, Qwen), large EP compute fusion operator optimization, key operator optimization (Attention MLA/GQA, Sparse Attention DSA, Linear Attention, GEMM, Group GEMM)

Responsibilities

Adapt mainstream and self-developed large models to multiple heterogeneous chips; resolve precision anomalies in adaptation; profile and optimize the full inference chain for performance; optimize inference framework/engine architecture for multi-chip characteristics; develop and tune core operators for specific chip micro-architectures.

Seniority

Senior, hands-on IC

Sourced via tencent · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.