大模型存储SRE工程师
Core
Operate and stabilize large-scale distributed storage products, integrating SRE practices with AI/LLM-driven automation for fault analysis and capacity management.
Role type
Senior SRE Engineer (Storage & AIOps)
Builds
Distributed storage systems (object/block/file/table) and automated SRE/AIOps platforms
Domain
Cloud Infrastructure / Distributed Storage / AIOps
Deliverable
production ML models | infrastructure
Required skills
Linux system administration, distributed storage architecture, shell scripting, Python, Golang, observability (SLI/SLO), incident management, capacity planning
Preferred skills
DevOps toolchain development, machine learning for anomaly detection, LLM application for log analysis and knowledge retrieval, soft-hardware collaboration
Technologies
Linux, Shell, Python, Golang, Kubernetes (implied by SRE/DevOps context), distributed storage systems
Responsibilities
Manage daily operations including releases, capacity planning, and 7x24 stability; Architect automated tools and platforms for fault and change management; Implement AIOps solutions for alert noise reduction and capacity forecasting; Lead BCP governance and risk mitigation; Optimize storage performance and reliability using data-driven insights
Seniority
Senior, hands-on IC