混元文本/多模态预训练框架研发工程师(深圳/北京/上海/杭州)
Core
Develop and optimize large-scale pre-training frameworks for NLP and multi-modal models, supporting efficient training on 10,000+ GPUs.
Role type
Senior IC machine-learning engineer (large-scale training framework)
Builds
Large-scale pre-training frameworks for text-to-image, text-to-video, and text-to-3D models
Domain
Artificial Intelligence / Large Language Models / Multi-modal AI
Deliverable
production ML models
Required skills
DeepSpeed, Megatron, 3D parallelism, ZeRO, Flash-Attn, CUDA optimization, ViT, SD, DiT model training
Preferred skills
Operator-level CUDA optimization, large model hyperparameter tuning, performance evaluation of large models
Technologies
CUDA, DeepSpeed, Megatron, ViT, SD, DiT
Responsibilities
Optimize large model training frameworks for single-task 10,000+ GPU scale; Design NLP and multi-modal model structures and validate training efficiency; Accelerate training performance for text-to-image/video/3D; Optimize low-precision and large-window training performance.