CareerPlanSign in

AI异构硬件训练优化专家-Seed

上海💼 Full-time🗓 2026-09-28

Core

Optimizing performance and stability for large-scale LLM training on custom AI accelerator clusters for ByteDance's Seed team.

Role type

Senior IC AI hardware training optimization engineer

Builds

High-throughput, stable training runs for proprietary LLMs supporting 50+ applications (e.g., Doubao, Jimeng)

Domain

AI Infrastructure / High-Performance Computing / Distributed Systems

Deliverable

production ML models

Required skills

C/C++ or Python, Linux, Computer Architecture, Chip Microarchitecture, High-Performance Computing, Distributed Systems, Parallel Computing

Preferred skills

GPU/NPU cluster optimization, NCCL/HCCL/RDMA/RoCE, MoE training, Mixed-precision training, ZeRO, Checkpoint optimization, Fault tolerance, Silent data corruption detection

Technologies

PyTorch, Megatron, DeepSpeed, FSDP, TorchTitan, NCCL, HCCL, RDMA, RoCE

Responsibilities

Optimize data, tensor, pipeline, and expert parallelism strategies; Resolve network jitter, communication bottlenecks, and silent data errors; Collaborate on compiler stacks and compute architectures for cluster stability.

Seniority

Senior, hands-on IC

Sourced via bytedance · Listed on CareerPlan, which tracks 878,000+ jobs from 20+ sources.