SRE专家/架构师(数据管理方向)-集团信息系统
Core
Establishing and managing an operations system for cost, quality, and efficiency; leading architectural evolution and system engineering for observability, maintainability, and high stability; solving service governance and compliance challenges in large-scale concurrency, high complexity, and AI scenarios.
Role type
Senior Site Reliability Engineer / Architect (Data Management)
Builds
Scalable, stable, and observable data management systems for enterprise-level applications
Domain
Internet / Cloud Computing / Data Management
Deliverable
production ML models | infrastructure
Required skills
Linux, networking, storage, Python, Go, full-stack troubleshooting, automation scripting, service governance, compliance
Preferred skills
Chaos engineering, AIOps, FinOps, fluent English
Technologies
Python, Go, Linux, Kubernetes (implied by SRE context), observability tools