大模型平台策略推理优化工程师(J97422)
Core
Design and implement strategies for cost reduction and efficiency optimization in large model inference for the Qianfan MaaS platform, including custom performance tuning.
Role type
Senior IC large model inference optimization engineer
Builds
Optimized inference systems and performance evaluation frameworks for the Qianfan MaaS platform
Domain
Large Language Models (LLMs), Inference Optimization, Cloud Infrastructure
Deliverable
production ML models
Required skills
Python, PyTorch, Transformer architecture, model quantization (PTQ/QAT), speculative decoding (MTP/Eagle), structured pruning, sparse models
Preferred skills
vLLM, SGLang, TensorRT-LLM, QAT, MTP training, Eagle/Medusa/MTP variants, large-scale online inference service optimization
Technologies
Python, PyTorch, vLLM, SGLang, TensorRT-LLM, INT4/INT8/FP8, PTQ, QAT, MTP, Eagle, Medusa
Responsibilities
Design and implement inference cost reduction and efficiency optimization strategies; Develop quantization and speculative decoding optimization schemes; Build inference performance evaluation and benefit assessment systems; Research and deploy frontier inference optimization technologies.
Seniority
Senior, hands-on IC