AI Infra研发工程师-上海/北京/深圳
Core
Evaluating GPU performance, developing training/inference frameworks, and optimizing large model inference for cloud customers.
Role type
Senior IC AI Infrastructure Engineer (GPU & Large Model Optimization)
Builds
Accelerated training and inference frameworks for public cloud customers
Domain
Cloud computing, AI infrastructure, Large Language Models (LLMs)
Deliverable
production ML models
Required skills
GPU architecture analysis, vLLM/SGlang/TensorRT-LLM, DeepSeek/Qwen/Flux models, Megatron-LM/DeepSpeed/verL, CUDA programming, NCCL/MPI, distributed parallelism, performance bottleneck analysis
Preferred skills
Autonomous driving algorithms (BEVformer/MapTRV2/SparseDrive/FlashOCC/Pointpillars), operator fusion, multi-node debugging
Responsibilities
Lead GPU performance benchmarking and adaptation for domestic and NVIDIA chips; Develop acceleration schemes for training and inference frameworks; Optimize large model inference performance for cost efficiency; Analyze and resolve performance bottlenecks in POCs and production environments; Translate academic research into framework features
Seniority
Senior, hands-on IC