CareerPlanSign in

Site Reliability Engineer

Headquarters💼 Full-time🗓 2026-07-27 → 2026-09-26

Core

Define reliability standards and build observability infrastructure for a K-12 workflow management platform from zero.

Role type

Senior IC Site Reliability Engineer (build-it-from-zero)

Builds

Production ML models | product features | infrastructure

Domain

Education technology (K-12 workflow management)

Deliverable

production ML models | infrastructure

Required skills

Systems literacy (OS, databases, networking), AI tooling for code/debugging, SLI/SLO methodology, Observability stack ownership, Incident management design, Chaos engineering, Infrastructure as Code, Cloud platforms, Python/Go/Bash

Preferred skills

None stated

Technologies

Grafana, Prometheus, Loki, OpenTelemetry, PagerDuty, Terraform, Kubernetes, AWS/GCP/Azure, Python, Go, Bash, k6, Locust, JMeter

Responsibilities

Define SLIs/SLOs and implement Grafana dashboards/alerts; Stand up and improve incident management practices; Own end-to-end observability stack (metrics, logs, traces, RUM); Partner with engineering on error budgets; Automate manual operational work; Design and run load/performance tests and chaos engineering

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.