Site Reliability Engineer - Infrastructure
Core
Managing complex infrastructure systems (Storage/Computing/DB) to ensure reliability, efficiency, and cost optimization at scale.
Role type
Senior Site Reliability Engineer (Infrastructure)
Builds
Automated operation solutions, monitoring frameworks, and high-availability architecture for large-scale systems.
Domain
Cloud Infrastructure & Big Data Systems
Deliverable
infrastructure
Required skills
Linux OS, Python, Go, Java, System Design, Capacity Planning, Bottleneck Analysis, Data Protection
Preferred skills
Kubernetes, Docker, Spark, Flink, Service Mesh, RPC Framework, AIops, KV/Graph/NoSQL DBs
Technologies
Kubernetes, Docker, Spark, Flink, Kafka, Redis, MySQL, MongoDB, MQ, RPC Framework, Service Mesh, AIops
Responsibilities
Ensure reliability and efficiency of core infrastructure; troubleshoot technical issues and manage high-availability architecture; build automated operation solutions; design monitoring frameworks for SOA governance; optimize system costs; design and implement data protection plans.
Seniority
Senior, hands-on IC