CareerPlanSign in

混元多模态大模型推理加速工程师(深圳/北京/上海/杭州)

Shenzhen, China💼 Full-time🗓 2026-09-28

Core

Deploy and optimize inference for general-purpose multimodal large models including understanding, generation, and speech models.

Role type

Senior IC multimodal large model inference acceleration engineer

Builds

Production multimodal models (understanding, generation, speech) and optimized inference frameworks

Domain

AI infrastructure, high-performance computing, multimodal large models

Deliverable

production ML models

Required skills

vllm, sglang, TensorRT, FasterTransformer, deep learning framework internals, operator fusion, quantization strategies, dynamic batching, KV cache optimization, distributed inference deployment

Preferred skills

industry deployment cases, open-source projects, model training/inference tuning, CPU/GPU acceleration, distributed training, top-tier academic publications in VQA, image generation, video understanding, speech recognition/synthesis

Technologies

vllm, sglang, TensorRT, FasterTransformer

Responsibilities

Deploy multimodal large models for understanding, generation, and speech; optimize inference framework performance and cost; track and implement cutting-edge technology; adapt deployment solutions for specific business needs

Seniority

Senior, hands-on IC

Sourced via tencent · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.