CareerPlanSign in

腾讯云-MaaS平台近线/离线推理研发专家

Shenzhen, China💼 Full-time🗓 2026-09-28

Core

Designing and optimizing offline (batch) inference platforms for large-scale LLM workloads to maximize GPU utilization and reduce costs.

Role type

Senior IC machine-learning infrastructure engineer (offline inference)

Builds

Scalable batch inference platforms with task scheduling, checkpointing, and priority queuing for LLMs.

Domain

Cloud computing / Large Language Model (LLM) inference infrastructure

Deliverable

production ML models

Required skills

Distributed systems architecture, LLM inference engine optimization (vLLM, SGLang, TensorRT-LLM), Quantization strategies, Resource scheduling, Cost modeling, SLA definition

Preferred skills

Experience building inference platforms from scratch, GPU resource management

Technologies

vLLM, SGLang, TensorRT-LLM

Responsibilities

Designing offline inference platform architecture including task scheduling and elastic scaling; Optimizing inference throughput and cost via engine selection and quantization; Managing resource sharing between online and offline inference; Building SLA systems and cost accounting models; Integrating with data synthesis and model evaluation pipelines.

Sourced via tencent · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.