CareerPlanSign in

AI大模型优化工程师 - 音视频技术

北京💼 Full-time🗓 2026-09-28

Core

Optimizing inference performance of self-developed and open-source LLM/VLM multimodal models across heterogeneous chip architectures to support massive-scale computing for video, live streaming, and offline data processing.

Role type

Senior IC AI model inference optimization engineer (multimodal/heterogeneous chips)

Builds

High-performance inference services for video, live streaming, and large-scale offline data processing using tens of thousands of GPU/NPU cards.

Domain

AI/ML inference optimization, heterogeneous computing, video streaming infrastructure

Deliverable

production ML models

Required skills

C/C++, Python, Linux, computer architecture, parallel computing, GPU/NPU hardware architecture, CUDA, CUTLASS, AscendC, BangC, LLM/VLM model structures, inference frameworks (vLLM, SGLang), model parallelism strategies, operator optimization, quantization algorithms

Preferred skills

Research on cutting-edge inference acceleration techniques, hardware-software co-optimization, new heterogeneous hardware characteristics

Technologies

CUDA, CUTLASS, AscendC, BangC, vLLM, SGLang, heterogeneous chips (GPU/NPU)

Responsibilities

Optimize inference of self-developed and open-source LLM/VLM models on multiple heterogeneous chips; Improve the adaptation and optimization efficiency of the inference technology stack for different chips; Analyze and evaluate new heterogeneous chips for various large models; Research and implement cutting-edge technologies such as inference acceleration and hardware-software co-optimization.

Seniority

Senior, hands-on IC

Sourced via bytedance · Listed on CareerPlan, which tracks 853,000+ jobs from 20+ sources.