机器学习平台SRE工程师-AML
Core
Ensure stability of machine learning systems supporting model development, training, and deployment; manage GPU/NPU/CPU and storage resources; handle multi-region disaster recovery and cluster governance.
Role type
Senior SRE for Machine Learning Platforms
Builds
Automated tools and platforms for resource management and operational efficiency
Domain
Cloud Infrastructure / Machine Learning Operations
Deliverable
infrastructure
Required skills
Linux administration, Go, Python, Shell scripting, Kubernetes, distributed system resource management, task scheduling, cluster governance, disaster recovery planning, cost optimization
Preferred skills
Large-scale distributed system operations, GPU server operations
Responsibilities
Manage and plan GPU/NPU/CPU and storage resources; develop automation tools to improve resource utilization and operational efficiency; oversee multi-region disaster recovery and service deployment; govern cluster machines.
Seniority
Senior, hands-on IC