大模型推理引擎专家 - Seed Model
Core
Design and develop large-scale machine learning system architecture for online and batch inference services, optimizing model performance and GPU cluster utilization for search, recommendation, and content moderation scenarios.
Role type
Senior IC machine learning systems engineer (inference engine)
Builds
High-concurrency, high-reliability inference frameworks and GPU cluster scheduling systems
Domain
AI infrastructure, large language models, GPU computing
Deliverable
production ML models
Required skills
C/C++, Python, Linux, PyTorch, TensorFlow, GPU programming, compiler optimization, model quantization, GPU cluster scheduling
Preferred skills
Recommendation/ad/search offline-online inference architecture, GPU hardware architecture, CUDA, cuDNN, performance analysis
Responsibilities
Design and develop large-scale ML system architecture for inference services; Provide high-performance model optimization solutions for ML frameworks; Manage elastic scheduling and GPU overcommitment for global GPU clusters; Collaborate with algorithm teams for joint optimization
Seniority
Senior, hands-on IC