大模型训练框架研发工程师 (训练效率优化方向)(J103616)
Core
Develop and optimize distributed training frameworks for large-scale foundation and specialized models to improve throughput, resource utilization, and stability across pre-training, SFT, RL, DPO, and knowledge distillation scenarios.
Role type
Senior IC machine-learning engineer (distributed training framework optimization)
Builds
High-performance distributed training systems for billion/trillion parameter models
Domain
AI/ML, Large Language Models, Distributed Systems
Deliverable
production ML models
Required skills
C/C++/Python, CUDA programming, GPU architecture, PyTorch internals, Megatron-LM/DeepSpeed/FSDP, NCCL optimization, MoE/Attention/ViT operator optimization, mixed-precision training, low-bit quantization
Preferred skills
DualPipe/async communication techniques, performance analysis tools (Nsight Systems/Compute, PyTorch Profiler), open-source contributions (MLSys/NeurIPS papers, Megatron-LM/DeepSpeed), inference engine knowledge (sglang, vLLM)
Technologies
CUDA, CUTLASS, Triton, TileLang, NCCL, PyTorch, BF16, FP8
Responsibilities
Optimize distributed training strategies (DP/TP/PP/CP/EP) for long-context and massive parameter models; design and implement custom high-performance operators for specific model structures; analyze and resolve stability issues like NCCL timeouts and OOM; integrate advanced acceleration techniques like communication compression and overlapping communication.
Seniority
Senior, hands-on IC