AIGC推理优化工程师
Core
Optimize inference efficiency for LLM, MultiModal-LLM, and T2I models by implementing operator optimization, quantization, pruning, and distillation to maximize GPU performance.
Role type
Senior IC AIGC inference optimization engineer
Builds
Low-latency, high-throughput, high-stability inference systems and frameworks for AIGC services
Domain
Artificial Intelligence / Large Language Models / GPU Computing
Deliverable
production ML models
Required skills
C++, Python, data structures and algorithms, concurrent programming, PyTorch, TensorFlow, PaddlePaddle, TensorRT-LLM, vLLM, model quantization, model pruning, model distillation
Preferred skills
experience with AIGC model training and inference optimization
Responsibilities
Optimize inference frameworks and deployment pipelines; research and implement new technologies to enhance inference performance; collaborate with business teams to align optimizations with requirements