元宝-LLM大模型推理工程师
Core
Develop and optimize large language model (LLM) inference frameworks and high-performance inference engines.
Role type
Senior IC LLM inference optimization engineer
Builds
High-performance LLM inference engines
Domain
AI/ML infrastructure, GPU computing
Deliverable
production ML models
Required skills
GPU high-performance computing optimization, CUDA programming, computer architecture, parallel computing, memory optimization, low-bit computation, deep learning framework internals (PyTorch/TensorFlow), LLM model acceleration techniques (subgraph matching, compilation, quantization)
Preferred skills
Experience with TensorRT-LLM, vLLM, AI engineering optimization
Technologies
CUDA, GPU, TensorRT-LLM, vLLM, PyTorch, TensorFlow
Responsibilities
Develop and optimize LLM inference frameworks; Optimize high-performance LLM inference engines using GPU and CUDA; Research and introduce forward-looking technical architectures for LLM training and inference.
Seniority
Senior, hands-on IC