机器学习系统SRE工程师 - Seed Model
Core
Maintaining stability of machine learning systems to support large model development, training, and deployment.
Role type
Senior Machine Learning Systems SRE
Builds
Stable GPU infrastructure and resource management for large model training and inference
Domain
AI / Large Language Models / Distributed Systems
Deliverable
infrastructure
Required skills
Linux administration, Go, Python, Shell, Kubernetes, Docker, distributed system resource management, task scheduling, cluster governance, multi-region disaster recovery, cost optimization
Preferred skills
None stated
Technologies
Kubernetes, Docker, Kata, Linux, Go, Python, Shell
Responsibilities
Maintain stability of ML systems supporting model dev/training/deploy; Manage and plan group GPU resources and costs; Govern cluster and service stability to improve resource utilization and operational efficiency; Manage multi-region/multi-datacenter disaster recovery and deployment; Provide operational support and troubleshoot system/business issues
Seniority
Mid-level, hands-on IC