微信-WeLM 大模型预训练框架研发工程师(深圳、上海)
Core
Build and optimize large-scale pre-training frameworks for multimodal data (text, audio, image) and post-training alignment processes.
Role type
Senior IC machine-learning engineer (LLM pre-training & alignment)
Builds
Production-scale pre-training and post-training frameworks for large language models
Domain
Artificial Intelligence / Large Language Models / Distributed Systems
Deliverable
production ML models
Required skills
Deep learning frameworks (PyTorch, JAX, TensorFlow), Transformer architecture, Distributed training optimization (context parallel, 2D/rotary attention, hybrid parallelism, activation recomputation), CUDA optimization, RL/RLHF, Reward modeling, Mixture of Experts, Continual pretraining
Preferred skills
Experience with Megatron or DeepSpeed, Experience in audio/visual model training
Technologies
PyTorch, JAX, TensorFlow, CUDA, Megatron, DeepSpeed
Responsibilities
Optimize distributed training for long sequences to improve throughput and cost-efficiency; Build post-training pipelines including RL, RLHF, and alignment; Collaborate with algorithm and data teams to automate the full workflow from data processing to deployment.
Seniority
Senior, hands-on IC