CareerPlanSign in

AI服务端架构师-剪映CapCut(北京/上海/深圳)

北京💼 Full-time🗓 2026-09-28

Core

Engineering deployment and inference optimization of key AI models for the CapCut product suite, including GPU resource management and utilization optimization.

Role type

Senior IC AI Infrastructure Engineer (Model Deployment & Optimization)

Builds

Inference engines and GPU management systems for CapCut's AI features

Domain

Consumer Media / AI Infrastructure

Deliverable

production ML models

Required skills

C++, Golang, CUDA, PyTorch, Diffusion/DiT model architecture, model quantization, pruning, distillation, concurrent programming, data structures and algorithms

Preferred skills

TensorRT-LLM, vLLM, xDiT, LightX2V, model training optimization

Technologies

CUDA, TensorRT, Cutlass, PyTorch, xDiT, LightX2V, TensorRT-LLM, vLLM

Responsibilities

Deploy and optimize inference for key CapCut AI models; Build and manage GPU resource systems across infrastructure platforms; Optimize GPU utilization for the CapCut product suite; Implement industry inference acceleration methods and tools.

Seniority

Senior, hands-on IC

Sourced via bytedance · Listed on CareerPlan, which tracks 846,000+ jobs from 20+ sources.