Staff Cloud Backend Engineer- Observability and Site Reliability
Core
Design and operate scalable observability platforms and SRE practices to ensure the reliability, performance, and availability of datacentre services.
Role type
Staff Cloud Backend Engineer (Observability and Site Reliability)
Builds
Large-scale Observability and Telemetry collection platforms, resilient datacentre infrastructure
Domain
E-commerce, Cloud Infrastructure, Datacentre Operations
Deliverable
production ML models | infrastructure
Required skills
Go or Python, Kubernetes internals, distributed systems, cloud-native architectures, load balancing, service mesh, root cause analysis, performance optimization
Preferred skills
LLM inference infrastructure, low-level optimization, hybrid-cloud management, regulated environment compliance
Technologies
Kubernetes, AWS, Azure, GCP
Responsibilities
Design and maintain observability solutions for datacentre infrastructure; Develop and optimize monitoring systems for high availability; Implement SRE best practices for reliability and scalability; Conduct root cause analysis and post-mortem reviews; Collaborate with engineering teams and vendors on observability requirements; Ensure compliance with security policies and industry standards
Seniority
Staff, hands-on IC