AI服务端架构师-剪映CapCut(北京/上海/深圳)
Core
Engineering deployment and inference optimization of key AI models for the CapCut product suite, including GPU resource management and utilization optimization.
Role type
Senior IC AI Infrastructure Engineer (Model Deployment & Optimization)
Builds
Inference engines and GPU management systems for CapCut's AI features
Domain
Consumer Media / AI Infrastructure
Deliverable
production ML models
Required skills
C++, Golang, CUDA, PyTorch, Diffusion/DiT model architecture, model quantization, pruning, distillation, concurrent programming, data structures and algorithms
Preferred skills
TensorRT-LLM, vLLM, xDiT, LightX2V, model training optimization
Technologies
CUDA, TensorRT, Cutlass, PyTorch, xDiT, LightX2V, TensorRT-LLM, vLLM
Responsibilities
Deploy and optimize inference for key CapCut AI models; Build and manage GPU resource systems across infrastructure platforms; Optimize GPU utilization for the CapCut product suite; Implement industry inference acceleration methods and tools.
Seniority
Senior, hands-on IC