SRE AI高级工程师-基础架构
Core
Manage and ensure consistency of massive high-performance GPU/XPU clusters for large model training, online inference, search, and recommendation training.
Role type
Senior SRE (AI Infrastructure)
Builds
Production GPU/XPU clusters for AI workloads
Domain
AI Infrastructure / High-Performance Computing
Deliverable
infrastructure
Required skills
GPU/XPU resource management, cluster scheduling, system architecture design, Go/Python/Java/C++, production troubleshooting, performance tuning, distributed training frameworks (TensorFlow/PyTorch)
Preferred skills
NVIDIA H100/A100 optimization, Ascend/XPU experience, global collaboration, project management
Technologies
NVIDIA H100, A100, Ascend, TensorFlow, PyTorch
Responsibilities
Deliver and guarantee consistency of massive GPU/XPU resources across training and inference scenarios; build automation to ensure stability and resource availability; isolate faulty GPU resources; design and manage the lifecycle of production clusters and services.