Site Reliability Engineer( SRE)
Core
Ensure stability, reliability, and high availability of overseas production environments through automation, observability, and incident response.
Role type
Site Reliability Engineer (SRE)
Builds
Automation platforms, engineering efficiency tools, and observability systems
Domain
Cloud infrastructure, distributed systems, and AI-driven operations
Deliverable
production ML models | infrastructure
Required skills
Python, Go, Java, C++, Linux, computer networking, load balancing, distributed systems, high-availability architectures, system monitoring, observability tools, scripting, CI/CD concepts
Preferred skills
Multi-cloud or hybrid cloud platforms, AI tools, AI Agents, AIOps technologies, disaster recovery planning, dependency management, traffic governance
Technologies
Prometheus, Grafana, ELK, Alibaba Cloud, Azure, AWS, GCP
Responsibilities
Manage resource provisioning, capacity planning, monitoring, change management, and incident response; Review system architecture and implement mitigations; Build and enhance observability systems; Develop and optimize automation platforms; Explore and promote AI technologies in operations scenarios; Collaborate with R&D, product, security, and infrastructure teams
Seniority
Mid-level, hands-on IC