CareerPlanSign in

微信-AI Infra工程师-大模型推理方向

Beijing, China💼 Full-time🗓 2026-09-28

Core

Design and optimize high-throughput, low-latency inference systems for large language models (LLMs) to support algorithm deployment and cost control.

Role type

Senior IC AI Infrastructure Engineer (LLM Inference)

Builds

Production-grade LLM inference frameworks and systems

Domain

Artificial Intelligence / Large Language Models / System Performance

Deliverable

production ML models

Required skills

C/C++, Python, GPU programming (CUDA, OpenCL), deep learning inference frameworks (TensorRT, vLLM, SGLang), distributed inference, heterogeneous CPU/GPU acceleration, model operator optimization

Preferred skills

Computer architecture background, server-side AI chip experience, large-scale model distributed deployment

Technologies

TensorRT, FasterTransformer, TensorRT-LLM, vLLM, SGLang, CUDA, OpenCL, cuBLAS, cuDNN, CUTLASS

Responsibilities

Collaborate with algorithm engineers to deploy deep learning algorithms; Optimize LLM inference performance for throughput and cost; Improve inference framework usability and debuggability.

Sourced via tencent · Listed on CareerPlan, which tracks 877,000+ jobs from 20+ sources.