CareerPlanSign in

混元大模型Infra稳定性研发工程师(深圳/北京/上海/杭州)

Shenzhen, China💼 Full-time🗓 2026-09-28

Core

Build and maintain stability for the Hunyuan large model infrastructure, focusing on fault detection, metric collection, and automated task recovery.

Role type

Senior IC infrastructure engineer (large-scale ML training)

Builds

Fault detection platforms, automated task retraining capabilities, and stability metrics for the Hunyuan model training pipeline.

Domain

Large-scale AI model training infrastructure

Deliverable

production ML models

Required skills

GPU/NPU architecture, RDMA networking, collective communication (all2all, allGather), PyTorch/Megatron training frameworks, containerization (Docker), large-scale system troubleshooting

Preferred skills

Experience troubleshooting large-scale task system failures

Technologies

PyTorch, Megatron, RDMA, Docker, GPU, NPU

Responsibilities

Establish stability governance and standards for the Hunyuan infrastructure, enhance metric collection across framework/compute/network modules, build platform capabilities for detecting faulty and slow nodes, collaborate on unified automatic task retraining, and resolve daily operational faults.

Seniority

Senior, hands-on IC

Sourced via tencent · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.