Staff Site Reliability Engineer
Core
Define and evolve SLOs, SLIs, and error budgets; lead incident management and blameless postmortems; design observability platforms; and build automation to reduce toil for KEV's mission-critical financial software.
Role type
Staff Site Reliability Engineer (Strategy & Hands-on IC)
Builds
Cloud-native financial software for K12 schools (payments, accounting, reporting)
Domain
Fintech / K12 Education / Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
SRE strategy, SLO/SLI/Error Budget definition, incident management, observability platform design, automation engineering, capacity planning, AI tooling evaluation, technical mentorship
Preferred skills
None stated
Technologies
Microsoft Azure, .NET, .NET Framework, IIS
Responsibilities
Define and evolve SLOs, SLIs, and error budgets; own incident management and drive blameless postmortems; design and evolve monitoring, alerting, logging, and tracing platforms; build automation tools to replace manual operational work; lead capacity planning and performance analysis; evaluate and apply AI tooling for reliability operations; coach and mentor engineers on reliability practices
Seniority
Staff, hands-on IC with strategic scope