硬件加速训练AI Infra工程师-Data
Core
Develop self-researched hardware training frameworks and optimize distributed training for large models on proprietary hardware.
Role type
Senior IC hardware-accelerated AI infrastructure engineer (training)
Builds
Self-researched training frameworks (Torch, Megatron, DTensor), distributed parallelism implementations, and optimized operators for compute/communication fusion.
Domain
AI Infrastructure / Large Language Models / Heterogeneous Computing
Deliverable
production ML models
Required skills
Heterogeneous hardware (GPU/NPU) training development, PyTorch, FSDP, Megatron-LM, VeRL, Distributed training patterns (DP/SP/TP/PP), Operator development, Compiler scheduling optimization, Computer architecture, Parallel computing.
Preferred skills
Large-scale cluster training (100+ cards), AI chip development and evaluation, FPGA development.
Responsibilities
Develop self-researched hardware training frameworks; Support training tasks for business large models on proprietary hardware; Develop and optimize distributed parallelism methods; Research and optimize training communication, computation, and compute-communication fusion operators.
Seniority
Mid-Senior, hands-on IC
