CareerPlanSign in

大模型平台策略推理优化工程师(J97422)

北京市💼 Full-time🗓 2026-07-21 → 2026-09-28

Core

Design and implement strategies for cost reduction and efficiency optimization in large model inference for the Qianfan MaaS platform, including custom performance tuning.

Role type

Senior IC large model inference optimization engineer

Builds

Optimized inference systems and performance evaluation frameworks for the Qianfan MaaS platform

Domain

Large Language Models (LLMs), Inference Optimization, Cloud Infrastructure

Deliverable

production ML models

Required skills

Python, PyTorch, Transformer architecture, model quantization (PTQ/QAT), speculative decoding (MTP/Eagle), structured pruning, sparse models

Preferred skills

vLLM, SGLang, TensorRT-LLM, QAT, MTP training, Eagle/Medusa/MTP variants, large-scale online inference service optimization

Technologies

Python, PyTorch, vLLM, SGLang, TensorRT-LLM, INT4/INT8/FP8, PTQ, QAT, MTP, Eagle, Medusa

Responsibilities

Design and implement inference cost reduction and efficiency optimization strategies; Develop quantization and speculative decoding optimization schemes; Build inference performance evaluation and benefit assessment systems; Research and deploy frontier inference optimization technologies.

Seniority

Senior, hands-on IC

Sourced via baidu · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.