CareerPlanSign in

高性能计算工程师

Shanghai, China💼 Full-time🗓 2026-09-28

Core

Design and optimize production-grade inference architectures for trillion-parameter LLMs, focusing on extreme performance, low-bit quantization, and heterogeneous hardware adaptation.

Role type

Senior IC high-performance computing engineer (LLM inference)

Builds

High-throughput LLM inference engines and distributed serving systems

Domain

AI infrastructure / Large Language Models / Heterogeneous Computing

Deliverable

production ML models

Required skills

C++, Python, CUDA/Triton, Transformer architecture, vLLM, TensorRT-LLM, low-bit quantization, distributed systems, kernel optimization, hardware ISA/microarchitecture

Preferred skills

Domestic chip adaptation (Ascend/Hygon), MoE scheduling, unified AI inference engine design

Technologies

vLLM, TensorRT-LLM, HCCL, NCCL, Docker, Kubernetes

Responsibilities

Architect and optimize core scheduling strategies like PagedAttention and continuous batching for trillion-parameter models; Implement industrial-grade low-bit quantization (INT4/FP8) and MoE distributed optimization; Design scalable unified AI inference engines for multimodal tasks; Optimize inference engines for domestic AI chips and customize communication primitives; Develop high-performance custom operators (Attention, GEMM, KV Cache) and leverage hardware ISA for peak performance.

Seniority

Senior, hands-on IC

Sourced via tencent · Listed on CareerPlan, which tracks 846,000+ jobs from 20+ sources.