模型系统工程师-抖音推荐
Core
Designing and optimizing large model inference systems and GPU resource management for Douyin's content understanding and review platforms.
Role type
Senior IC machine-learning systems engineer (large model inference)
Builds
Large model inference frameworks, RAG applications, and GPU/CPU resource governance systems
Domain
Internet / Large Language Models / Distributed Systems
Deliverable
production ML models
Required skills
C/C++ or Python, CUDA programming, distributed system design, large model architecture knowledge, model quantization, TRT-LLM/vLLM, kernel optimization
Preferred skills
ACM/ICPC/NOI/IOI competition awards, experience with model pruning, large-scale distributed system maintenance
Technologies
TensorFlow, PyTorch, TRT-LLM, vLLM, CUDA
Responsibilities
Design and optimize large model inference system architecture, research and implement inference acceleration techniques (quantization, distillation, kernel optimization), manage and govern GPU resources to improve efficiency
Seniority
Senior, hands-on IC