CareerPlanSign in

AI Infra研发工程师-上海/北京/深圳

Shanghai, China💼 Full-time🗓 2026-09-28

Core

Evaluating GPU performance, developing training/inference frameworks, and optimizing large model inference for cloud customers.

Role type

Senior IC AI Infrastructure Engineer (GPU & Large Model Optimization)

Builds

Accelerated training and inference frameworks for public cloud customers

Domain

Cloud computing, AI infrastructure, Large Language Models (LLMs)

Deliverable

production ML models

Required skills

GPU architecture analysis, vLLM/SGlang/TensorRT-LLM, DeepSeek/Qwen/Flux models, Megatron-LM/DeepSpeed/verL, CUDA programming, NCCL/MPI, distributed parallelism, performance bottleneck analysis

Preferred skills

Autonomous driving algorithms (BEVformer/MapTRV2/SparseDrive/FlashOCC/Pointpillars), operator fusion, multi-node debugging

Responsibilities

Lead GPU performance benchmarking and adaptation for domestic and NVIDIA chips; Develop acceleration schemes for training and inference frameworks; Optimize large model inference performance for cost efficiency; Analyze and resolve performance bottlenecks in POCs and production environments; Translate academic research into framework features

Seniority

Senior, hands-on IC

Sourced via tencent · Listed on CareerPlan, which tracks 846,000+ jobs from 20+ sources.