Senior/Lead Site Reliability Engineer
Core
Manage operational work for NetEase Interactive Entertainment services (e.g., Eggy Party, Marvel Rivals) and internal research projects using software engineering methods to achieve operational automation and improve service availability.
Role type
Senior/Lead Site Reliability Engineer
Builds
High-quality and efficient operational services for game servers at controllable costs
Domain
Gaming industry + Cloud infrastructure & DevOps
Deliverable
production ML models | infrastructure
Required skills
Linux administration, TCP/IP and HTTP protocols, C/C++/Shell/Python/Golang/Rust/Java, open-source software knowledge (Linux, Nginx, MySQL, K8S, Istio), AIOps analysis, automated script generation, root-cause analysis
Preferred skills
AI/LLM trends awareness, hands-on experience applying AI to operations, open-source community contributions
Technologies
Linux, Nginx, MySQL, K8S, Istio, C/C++, Shell, Python, Golang, Rust, Java
Responsibilities
Design and select basic runtime environments (servers, virtualization, cloud services, networks, databases) based on game service architecture and performance requirements; Establish and monitor various operational metrics and customize data analysis standards; Collaborate with product departments to identify issues, optimize technical architecture, and enhance user experience; Participate in in-depth research on cutting-edge open-source software, virtualization, databases, and web services to develop technical solutions
Seniority
Senior/Lead, hands-on IC