AML机器学习平台SRE工程师-Data
Core
Ensure stability of machine learning systems supporting model development, training, and deployment; manage GPU/NPU/CPU and storage resources; handle multi-region disaster recovery and cluster governance.
Role type
Senior Site Reliability Engineer (ML Infrastructure)
Builds
ML training and inference platforms, resource management systems, automation tools
Domain
Cloud infrastructure, Machine Learning, Distributed Systems
Deliverable
infrastructure
Required skills
Linux administration, Go, Python, Shell scripting, Kubernetes, Distributed system operations, Resource scheduling, Cluster governance, Disaster recovery planning, Cost optimization
Preferred skills
Large-scale distributed system operations, GPU server management
Technologies
Kubernetes, Linux, Go, Python, Shell, GPU/NPU/CPU, Storage systems
Responsibilities
Manage and plan GPU/NPU/CPU and storage resources including cost and budget; Implement multi-region disaster recovery and service deployment; Develop automation tools to improve resource utilization and operational efficiency; Maintain cluster machine governance.