CareerPlanSign in

大模型训练框架研发工程师 (训练效率优化方向)(J103616)

北京市💼 Full-time🗓 2026-07-28 → 2026-09-28

Core

Develop and optimize distributed training frameworks for large-scale foundation and specialized models to improve throughput, resource utilization, and stability across pre-training, SFT, RL, DPO, and knowledge distillation scenarios.

Role type

Senior IC machine-learning engineer (distributed training framework optimization)

Builds

High-performance distributed training systems for billion/trillion parameter models

Domain

AI/ML, Large Language Models, Distributed Systems

Deliverable

production ML models

Required skills

C/C++/Python, CUDA programming, GPU architecture, PyTorch internals, Megatron-LM/DeepSpeed/FSDP, NCCL optimization, MoE/Attention/ViT operator optimization, mixed-precision training, low-bit quantization

Preferred skills

DualPipe/async communication techniques, performance analysis tools (Nsight Systems/Compute, PyTorch Profiler), open-source contributions (MLSys/NeurIPS papers, Megatron-LM/DeepSpeed), inference engine knowledge (sglang, vLLM)

Technologies

CUDA, CUTLASS, Triton, TileLang, NCCL, PyTorch, BF16, FP8

Responsibilities

Optimize distributed training strategies (DP/TP/PP/CP/EP) for long-context and massive parameter models; design and implement custom high-performance operators for specific model structures; analyze and resolve stability issues like NCCL timeouts and OOM; integrate advanced acceleration techniques like communication compression and overlapping communication.

Seniority

Senior, hands-on IC

Sourced via baidu · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.