AML机器学习系统SRE工程师
Core
Maintaining stability of machine learning systems, supporting model development, training, and deployment, managing resources, and ensuring disaster recovery across multi-region clusters.
Role type
Senior SRE for Applied Machine Learning systems
Builds
High-performance, scalable ML system architecture and end-to-end ML services
Domain
Machine Learning Infrastructure / Distributed Systems
Deliverable
infrastructure
Required skills
Linux administration, Go, Python, Shell, Kubernetes, Docker, Kata, resource management, cost optimization, disaster recovery, cluster governance
Preferred skills
Large-scale distributed system operations, GPU server operations
Technologies
Kubernetes, Docker, Kata, Linux, Go, Python, Shell
Responsibilities
Maintain stability of ML systems across model development, training, and deployment; Manage and plan GPU/CPU resources and storage; Handle multi-region disaster recovery and cluster governance; Improve system stability, resource utilization, and operational efficiency.