Site Reliability Engineer
Core
Define reliability standards and build observability infrastructure for a K-12 workflow management platform from zero.
Role type
Senior IC Site Reliability Engineer (build-it-from-zero)
Builds
Production ML models | product features | infrastructure
Domain
Education technology (K-12 workflow management)
Deliverable
production ML models | infrastructure
Required skills
Systems literacy (OS, databases, networking), AI tooling for code/debugging, SLI/SLO methodology, Observability stack ownership, Incident management design, Chaos engineering, Infrastructure as Code, Cloud platforms, Python/Go/Bash
Preferred skills
None stated
Technologies
Grafana, Prometheus, Loki, OpenTelemetry, PagerDuty, Terraform, Kubernetes, AWS/GCP/Azure, Python, Go, Bash, k6, Locust, JMeter
Responsibilities
Define SLIs/SLOs and implement Grafana dashboards/alerts; Stand up and improve incident management practices; Own end-to-end observability stack (metrics, logs, traces, RUM); Partner with engineering on error budgets; Automate manual operational work; Design and run load/performance tests and chaos engineering
Seniority
Senior, hands-on IC