Software Engineer - Site Reliability Engineering
Core
Build systems, tools, and practices to improve the reliability and resilience of Neo4j Aura, a global DBaaS platform running on Kubernetes across major cloud providers.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Internal automation tools, observability stacks, and incident response processes for a large-scale distributed database service.
Domain
Cloud infrastructure, distributed systems, and database operations.
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Go, Python, Kubernetes, Terraform, Kustomize, CI/CD, observability, incident response, SRE principles, SLI/SLO definition, cloud architecture.
Preferred skills
Cluster-level administration, experience with OTel Collector, Prometheus, Grafana, Google Cloud operations suite.
Technologies
Go, Python, Kubernetes, Terraform, Kustomize, GitHub Actions, OTel Collector, Prometheus, Grafana, Google Cloud
Responsibilities
Build automation tools for troubleshooting and safe rollouts; define and act on SLIs and SLOs; shape the observability stack; participate in on-call rotations and incident response; write postmortems leading to lasting changes.
Seniority
Senior, hands-on IC
