Senior Staff Cloud Backend Engineer - Observability and Site Reliability
Core
Design, build, and operate scalable observability and reliability solutions for large-scale datacenter infrastructure, focusing on high-performance monitoring, telemetry, and automation.
Role type
Senior Staff IC Site Reliability Engineer (Data Centre Observability)
Builds
Large-scale observability and telemetry platforms, automation for infrastructure provisioning, and monitoring systems.
Domain
E-commerce / Cloud Infrastructure / Data Centre Operations
Deliverable
production ML models | infrastructure
Required skills
Go, Python, distributed systems, cloud-native architectures, Kubernetes, networked systems, performance optimization, root cause analysis, automation, SRE best practices
Preferred skills
None stated
Technologies
Kubernetes, Go, Python
Responsibilities
Design and implement observability solutions (monitoring, logging, alerting, telemetry); Develop and operate large-scale telemetry platforms; Lead root cause analysis and post-incident reviews; Analyze system performance to identify bottlenecks; Partner with cross-functional teams on reliability requirements; Ensure solutions adhere to security policies.
Seniority
Senior Staff, hands-on IC