Senior Systems Reliability Engineer
Core
Design, build, and operate software systems that make production applications measurable, scalable, and self-healing, treating operational challenges as engineering problems.
Role type
Senior Systems Reliability Engineer (Software Engineering)
Builds
Production reliability tooling, automation frameworks, observability stacks, and self-healing infrastructure components.
Domain
Energy grid operations / Software Engineering / Reliability Engineering
Deliverable
production ML models | product features | infrastructure
Required skills
Software engineering, SLO and error budget management, observability architecture, incident response leadership, NERC/CIP compliance, chaos engineering, capacity planning, distributed tracing, root cause analysis, Java/Spring Boot expertise
Preferred skills
None explicitly stated
Technologies
Grafana LGTM stack (Loki, Grafana, Tempo, Mimir), Dynatrace APM, Splunk, Datadog, Java, Spring Boot
Responsibilities
Design and build reliability tooling and automation frameworks; own SLO governance and error budget management; lead high-severity incident response and post-mortems; architect MLTP observability solutions; hold NERC/CIP compliance responsibility; participate in 24/7 on-call rotation; define and enforce production readiness standards.
Seniority
Senior, hands-on IC