SRE AI高级工程师-基础架构
Core
Deliver and ensure consistency of massive high-performance GPU/XPU resources for large-scale model training, online inference, search, and recommendation training clusters.
Role type
Senior SRE (AI Infrastructure)
Builds
Production-scale GPU/XPU clusters for AI workloads
Domain
AI Infrastructure / High-Performance Computing
Deliverable
infrastructure
Required skills
GPU/XPU resource management and scheduling, large-scale HPC cluster operations, system architecture design, Go/Python/Java/C++ development, production troubleshooting, performance tuning, distributed training frameworks (TensorFlow/PyTorch)
Preferred skills
Experience with NVIDIA H100/A100/Ascend/XPU, global collaboration, product and engineering mindset, project management, data structures and system design
Technologies
NVIDIA H100, A100, Ascend, TensorFlow, PyTorch, Go, Python, Java, C++