微信-AI Infra工程师-大模型推理方向
Core
Design and optimize high-throughput, low-latency inference systems for large language models (LLMs) to support algorithm deployment and cost control.
Role type
Senior IC AI Infrastructure Engineer (LLM Inference)
Builds
Production-grade LLM inference frameworks and systems
Domain
Artificial Intelligence / Large Language Models / System Performance
Deliverable
production ML models
Required skills
C/C++, Python, GPU programming (CUDA, OpenCL), deep learning inference frameworks (TensorRT, vLLM, SGLang), distributed inference, heterogeneous CPU/GPU acceleration, model operator optimization
Preferred skills
Computer architecture background, server-side AI chip experience, large-scale model distributed deployment
Technologies
TensorRT, FasterTransformer, TensorRT-LLM, vLLM, SGLang, CUDA, OpenCL, cuBLAS, cuDNN, CUTLASS
Responsibilities
Collaborate with algorithm engineers to deploy deep learning algorithms; Optimize LLM inference performance for throughput and cost; Improve inference framework usability and debuggability.