CareerPlanSign in

大模型训练框架开发工程师-抖音研发

北京💼 Full-time🗓 2026-09-28

Core

Design and optimize distributed training architectures for large language models, focusing on performance, resource utilization, and reliability across GPU/NPU platforms.

Role type

Senior IC distributed training framework engineer (LLM)

Builds

High-reliability training pipelines and toolchains for large-scale model training

Domain

Artificial Intelligence / Large Language Model Training

Deliverable

production ML models

Required skills

Distributed training strategies (FSDP, ZeRO, TP, PP, EP, SP), GPU/NPU performance tuning, multi-node multi-GPU communication mechanisms (NCCL, HCCL), PyTorch/DeepSpeed/Megatron/MindSpore framework internals, system analysis, CUDA/C++ optimization

Preferred skills

Experience training models >32B Dense or >100B MoE, contributions to open-source frameworks (Megatron-LM, DeepSpeed, ColossalAI, MindSpore), large-scale cluster training (100+ cards)

Responsibilities

Design and optimize distributed training architectures; Improve GPU/NPU training performance and resource utilization; Build high-reliability training pipelines (Checkpoint, mixed precision, fault tolerance, profiling); Adapt models to different hardware platforms; Provide training support and toolchain development for model teams

Sourced via bytedance · Listed on CareerPlan, which tracks 850,000+ jobs from 20+ sources.