混元大模型Infra稳定性研发工程师(深圳/北京/上海/杭州)
Core
Build and maintain stability for the Hunyuan large model infrastructure, focusing on fault detection, metric collection, and automated task recovery.
Role type
Senior IC infrastructure engineer (large-scale ML training)
Builds
Fault detection platforms, automated task retraining capabilities, and stability metrics for the Hunyuan model training pipeline.
Domain
Large-scale AI model training infrastructure
Deliverable
production ML models
Required skills
GPU/NPU architecture, RDMA networking, collective communication (all2all, allGather), PyTorch/Megatron training frameworks, containerization (Docker), large-scale system troubleshooting
Preferred skills
Experience troubleshooting large-scale task system failures
Technologies
PyTorch, Megatron, RDMA, Docker, GPU, NPU
Responsibilities
Establish stability governance and standards for the Hunyuan infrastructure, enhance metric collection across framework/compute/network modules, build platform capabilities for detecting faulty and slow nodes, collaborate on unified automatic task retraining, and resolve daily operational faults.
Seniority
Senior, hands-on IC