北京-SRE工程师(J100710)
Core
Ensure reliable, stable, and efficient operation of Baidu's large-scale distributed systems and online services.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Automated reliability systems, data center infrastructure, and large-scale traffic access solutions.
Domain
Internet / Distributed Systems / Cloud Infrastructure
Deliverable
production ML models
Required skills
Linux, C/C++/Python/Go/Shell, distributed system design, TCP/IP/HTTP/HTTPS, Socket programming, fault tracing, capacity management, elastic computing, AI for operations
Preferred skills
Docker, containerization, industry technology trends
Technologies
Docker, Linux, C, C++, Python, Go, Shell, TCP/IP, HTTP/HTTPS
Responsibilities
Design and develop service O&M solutions including web acceleration, continuous delivery, and performance tuning; Optimize large-scale traffic access systems and explore new technologies; Lead implementation of automated reliability systems.
Seniority
Senior, hands-on IC