SRE AI高级工程师-基础架构
Core
Deliver and ensure consistency of massive high-performance GPU/XPU resources for large-scale model training, online inference, search, and recommendation training clusters.
Role type
Senior Site Reliability Engineer (AI Infrastructure)
Builds
Production-scale GPU/XPU clusters for AI workloads
Domain
AI Infrastructure / High-Performance Computing
Deliverable
infrastructure
Required skills
GPU/XPU resource management and scheduling, large-scale HPC cluster operations, system architecture design, Go/Python/Java/C++ development, production troubleshooting, performance tuning, distributed training frameworks (TensorFlow/PyTorch)
Preferred skills
NVIDIA H100/A100 optimization, systematic engineering thinking, product and engineering mindset, data structure and system design
Technologies
NVIDIA H100, NVIDIA A100, Ascend, TensorFlow, PyTorch
Responsibilities
Manage clusters for multi-scenario AI workloads, build automation for stability and resource efficiency, design and deploy production cluster services, isolate and resolve GPU faults
Seniority
Senior, hands-on IC