SIte Reliability Engineer
Core
Ensure world-class resilience and performance for school management tools used by 7,000+ schools and trusts.
Role type
Site Reliability Engineer (SRE)
Builds
High-availability, scalable platform for school MIS and management tools
Domain
EdTech / Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
SRE practices, performance monitoring, capacity planning, scripting/automation, Infrastructure as Code (Terraform), relational databases (AWS Aurora), messaging/distributed workloads, nginx
Preferred skills
Enterprise scale solutions, containerization (Docker), Agile/Kanban, clean code, TDD
Technologies
DataDog, Prometheus, Terraform, AWS Aurora, nginx, Docker
Responsibilities
Monitor and analyze platform performance; collaborate on scalability and SLOs; improve observability with dashboards/alerts; devise runbooks and test DR/H/A plans; assess capacity and plan scaling; respond to incidents and lead blameless postmortems; maintain playbooks and documentation.
Seniority
Mid-Senior, hands-on IC