高级SRE运维工程师(资源成本方向)-抖音
Core
Design and implement automated tools and systems to manage, monitor, and adjust server resource usage for Douyin, focusing on capacity planning, cost reduction, and efficiency optimization.
Role type
Senior Site Reliability Engineer (Resource Cost & Capacity)
Builds
Automated resource management systems, capacity planning data systems, and auto-scaling mechanisms.
Domain
Internet / Cloud Infrastructure / Resource Management
Required skills
Python, Go, Java, C/C++, Cloud computing, Container technology, Microservices architecture, Load prediction, Cost optimization, Automation engineering
Responsibilities
Develop automated tools for resource management and monitoring; Analyze historical load and user behavior to predict future resource needs; Optimize resource efficiency and reduce waste; Support large-scale event resource planning; Collaborate with development and product teams for full lifecycle resource management.