机器学习系统SRE工程师 - Seed Model
Core
Maintaining stability of machine learning systems supporting large model development, training, and deployment.
Role type
Senior Machine Learning Systems SRE
Builds
Stable GPU infrastructure and resource management for large model training and deployment
Domain
AI / Large Language Models / Distributed Systems
Deliverable
infrastructure
Required skills
Linux administration, Go, Python, Shell, Kubernetes, Docker, distributed system resource management, task scheduling, cluster governance, multi-region disaster recovery
Preferred skills
Cost optimization, platform engineering, system stability governance
Responsibilities
Maintain stability of ML systems supporting large model dev/training/deploy; Manage and plan group GPU resources and costs; Govern cluster and service stability to improve resource utilization; Manage multi-region/multi-datacenter disaster recovery and deployment; Provide operational support and troubleshoot system issues
