AI服务端资深架构师-剪映CapCut(北京/上海/深圳)
Core
Engineering deployment and inference optimization of key AI models for the Jianying/CapCut product suite, including GPU resource management and utilization optimization.
Role type
Senior IC AI Infrastructure Engineer (AIGC/LLM)
Builds
Inference-optimized AI models and GPU management systems for video editing products
Domain
Consumer media + AI Infrastructure
Deliverable
production ML models
Required skills
C++, Golang, CUDA, PyTorch, Diffusion/DiT model architecture, model quantization, pruning, distillation, concurrent programming, data structures and algorithms
Preferred skills
TensorRT-LLM, vLLM, xDiT, LightX2V, model training optimization
Technologies
CUDA, TensorRT, Cutlass, PyTorch, xDiT, LightX2V, TensorRT-LLM, vLLM
Responsibilities
Deploy and optimize inference for key AI models; Build and manage GPU resource systems across infrastructure platforms; Optimize GPU utilization for the product suite; Implement inference acceleration techniques for AIGC models.
Seniority
Senior, hands-on IC