CareerPlanSign in

大模型训练稳定性和容错系统专家 - Seed Model

北京💼 Full-time🗓 2026-09-28

Core

Designing and optimizing stability architecture and fault-tolerance mechanisms for ultra-large-scale distributed training clusters to ensure reliable model training.

Role type

Senior IC distributed training stability and fault-tolerance engineer

Builds

Stable, fault-tolerant large-scale distributed training infrastructure for foundation models

Domain

AI / Large Language Models / Distributed Systems

Deliverable

production ML models

Required skills

Distributed training principles, Python/C++/Go, PyTorch, NCCL, RDMA, fault-tolerance mechanisms, automated root cause analysis, cluster stability metrics

Preferred skills

Megatron-LM, DeepSpeed, GPU hardware characteristics, domestic heterogeneous computing, training performance optimization

Technologies

PyTorch, Megatron-LM, DeepSpeed, NCCL, RDMA

Responsibilities

Design and iterate stability architecture for distributed training clusters; Develop and implement fault-tolerance mechanisms with second-level anomaly detection and automatic recovery; Build intelligent root cause analysis systems for training failures; Research and apply frontier technologies in large-scale training stability and optimization.

Sourced via bytedance · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.