混元多模态大模型推理加速工程师(深圳/北京/上海/杭州)
Core
Deploy and optimize inference for general-purpose multimodal large models including understanding, generation, and speech models.
Role type
Senior IC multimodal large model inference acceleration engineer
Builds
Production multimodal models (understanding, generation, speech) and optimized inference frameworks
Domain
AI infrastructure, high-performance computing, multimodal large models
Deliverable
production ML models
Required skills
vllm, sglang, TensorRT, FasterTransformer, deep learning framework internals, operator fusion, quantization strategies, dynamic batching, KV cache optimization, distributed inference deployment
Preferred skills
industry deployment cases, open-source projects, model training/inference tuning, CPU/GPU acceleration, distributed training, top-tier academic publications in VQA, image generation, video understanding, speech recognition/synthesis
Technologies
vllm, sglang, TensorRT, FasterTransformer
Responsibilities
Deploy multimodal large models for understanding, generation, and speech; optimize inference framework performance and cost; track and implement cutting-edge technology; adapt deployment solutions for specific business needs
Seniority
Senior, hands-on IC