AI异构硬件训练优化专家 - Seed Model
Core
Optimize performance and stability for large-scale LLM training on custom AI accelerator clusters, focusing on training frameworks, communication links, operator libraries, and cluster scheduling.
Role type
Senior IC AI hardware training optimization engineer
Builds
High-throughput, stable training runs for Seed's self-developed LLM models (e.g., Doubao) on multi-tenka GPU/NPU clusters
Domain
AI Infrastructure / High-Performance Computing / Distributed Systems
Deliverable
production ML models
Required skills
C/C++ or Python, Linux, Computer Architecture, Chip Microarchitecture, High-Performance Computing, Distributed Systems, Parallel Computing
Preferred skills
NCCL, HCCL, RDMA, RoCE, Collective Communication, Communication-Computation Fusion, MoE training, Mixed Precision training, ZeRO, Checkpoint optimization, Fault tolerance
Technologies
PyTorch, Megatron, DeepSpeed, FSDP, TorchTitan
Responsibilities
Optimize data, tensor, pipeline, and expert parallelism strategies to improve end-to-end throughput and scaling efficiency; Resolve silent data corruption, network jitter, communication bottlenecks, and exception localization; Drive software-hardware co-optimization across training frameworks, operator libraries, communication networks, compiler stacks, and compute architectures.
Seniority
Senior, hands-on IC
