Lead Site Reliability Engineer - Imunify Reliability Platform (remote work)
Core
Define and build a telemetry, alerting, and escalation system for ~70 Linux server security components to detect failures in hours instead of months.
Role type
Lead Site Reliability Engineer (SLO/Telemetry Platform)
Builds
A coherent SLO framework, collection pipeline, and alerting/escalation platform for a fleet of customer servers.
Domain
Linux server security infrastructure (WAF, IDS/IPS, malware scanning) and distributed systems telemetry.
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
SLO/SLI framework definition, Python, Go/Rust, time-series telemetry (Prometheus/Grafana), distributed systems debugging on bare metal, configuration management (Ansible), high-cardinality data handling.
Preferred skills
Security product background (WAF/EDR), audit monitoring (SOC 2/ISO 27001), OpenTelemetry, eBPF, agentic development tooling.
Technologies
Python, Go, Rust, Prometheus, Grafana, Alertmanager, ClickHouse, Ansible, GitLab CI, Jenkins, OpenTelemetry, eBPF.
Responsibilities
Define SLI taxonomy and SLOs for ~70 components, build telemetry collection pipelines, design SLO-anchored alerting and escalation systems, establish incident command practices and postmortems.
Seniority
Senior, hands-on IC with team leadership responsibilities.