高性能计算工程师
Core
Design and optimize production-grade inference architectures for trillion-parameter LLMs, focusing on extreme performance, low-bit quantization, and heterogeneous hardware adaptation.
Role type
Senior IC high-performance computing engineer (LLM inference)
Builds
High-throughput LLM inference engines and distributed serving systems
Domain
AI infrastructure / Large Language Models / Heterogeneous Computing
Deliverable
production ML models
Required skills
C++, Python, CUDA/Triton, Transformer architecture, vLLM, TensorRT-LLM, low-bit quantization, distributed systems, kernel optimization, hardware ISA/microarchitecture
Preferred skills
Domestic chip adaptation (Ascend/Hygon), MoE scheduling, unified AI inference engine design
Technologies
vLLM, TensorRT-LLM, HCCL, NCCL, Docker, Kubernetes
Responsibilities
Architect and optimize core scheduling strategies like PagedAttention and continuous batching for trillion-parameter models; Implement industrial-grade low-bit quantization (INT4/FP8) and MoE distributed optimization; Design scalable unified AI inference engines for multimodal tasks; Optimize inference engines for domestic AI chips and customize communication primitives; Develop high-performance custom operators (Attention, GEMM, KV Cache) and leverage hardware ISA for peak performance.
Seniority
Senior, hands-on IC