Site Reliability Engineer
Core
Build and scale a managed distributed database service across major cloud providers, focusing on infrastructure automation, observability, and reliability.
Role type
Junior Site Reliability Engineer (Cloud Infrastructure)
Builds
Managed database service on Kubernetes across AWS, Azure, and Google Cloud
Domain
Cloud Infrastructure / Distributed Databases
Deliverable
infrastructure
Required skills
Infrastructure automation, Python scripting, Bash scripting, Prometheus monitoring stack, Kubernetes, Cloud platforms (AWS/Azure/GCP), Production debugging
Preferred skills
Grafana, Mimir, Loki
Technologies
Kubernetes, Prometheus, Grafana, Mimir, Loki, AWS, Azure, Google Cloud
Responsibilities
Develop automation platform for infrastructure rollouts, Optimize telemetry platform for customer-impacting events, Partner with engineering to optimize cloud architecture performance, Debug live site events and conduct postmortem analysis, Participate in SLA-driven on-call rotation
Seniority
Junior (0-2 years experience)
