云计算运维研发Devops工程师(J94372)
Core
Design and implement scalable DevOps and SRE solutions for Baidu's large-scale distributed systems and cloud services, focusing on reliability, stability, and efficiency.
Role type
Senior Site Reliability Engineer (SRE) / DevOps Engineer
Builds
Automated platforms for service availability, large model training infrastructure, server lifecycle management, high-performance storage, and fault prediction systems.
Domain
Cloud Computing / Distributed Systems / AI Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Linux OS expertise, Python/Go/Shell programming, Cloud architecture design, Distributed systems knowledge, High-performance computing, Virtualization (KVM/OpenStack), Containerization (k8s), Fault prediction, Cost management
Preferred skills
Large model framework experience, High-performance communication protocols, OS/Kernel tuning, Industry-leading technology trends
Responsibilities
Lead architecture design for cloud systems, Develop automated platforms for service availability, Design technical solutions for large-scale operations, Monitor and optimize system performance and costs
Seniority
Mid-level, hands-on IC