CareerPlanSign in

微信 -WeLM 大模型推理优化工程师(深圳、上海)

Beijing, China💼 Full-time🗓 2026-09-28

Core

Optimize inference performance for large language models (LLMs) by reducing latency, increasing throughput, and improving resource efficiency across various hardware platforms.

Role type

Senior IC machine-learning inference optimization engineer

Builds

High-performance, scalable LLM inference services supporting real-time and batch processing scenarios

Domain

Artificial Intelligence / Large Language Models / GPU Computing

Deliverable

production ML models

Required skills

C++, Python, CUDA programming, GPU performance tuning, PyTorch, JAX, model quantization, model sparsification, Transformer architecture, KV Cache optimization, profiling tools (Nsight)

Preferred skills

Experience with dedicated AI chips, designing inference-friendly model architectures

Technologies

PyTorch, JAX, CUDA, Nsight, GPU, AI chips

Responsibilities

Develop and apply model compression techniques; optimize inference frameworks for multiple hardware platforms; design stable and scalable inference service architectures; establish performance benchmarking frameworks; analyze performance bottlenecks and implement optimization strategies; translate cutting-edge research into production environments.

Seniority

Senior, hands-on IC

Sourced via tencent · Listed on CareerPlan, which tracks 846,000+ jobs from 20+ sources.