CareerPlanSign in

Lead Site Reliability Engineer - Imunify Reliability Platform (remote work)

Sofia, Sofia City Province, Bulgaria🌐 Remote💼 Full-time🗓 2026-08-21 → 2026-09-25

Core

Define and build a telemetry, alerting, and escalation system for ~70 Linux server security components to detect failures in hours instead of months.

Role type

Lead Site Reliability Engineer (SLO/Telemetry Platform)

Builds

A coherent SLO framework, collection pipeline, and alerting/escalation platform for a fleet of customer servers.

Domain

Linux server security infrastructure (WAF, IDS/IPS, malware scanning) and distributed systems telemetry.

Deliverable

production ML models | product features | dashboards & analysis | infrastructure

Required skills

SLO/SLI framework definition, Python, Go/Rust, time-series telemetry (Prometheus/Grafana), distributed systems debugging on bare metal, configuration management (Ansible), high-cardinality data handling.

Preferred skills

Security product background (WAF/EDR), audit monitoring (SOC 2/ISO 27001), OpenTelemetry, eBPF, agentic development tooling.

Technologies

Python, Go, Rust, Prometheus, Grafana, Alertmanager, ClickHouse, Ansible, GitLab CI, Jenkins, OpenTelemetry, eBPF.

Responsibilities

Define SLI taxonomy and SLOs for ~70 components, build telemetry collection pipelines, design SLO-anchored alerting and escalation systems, establish incident command practices and postmortems.

Seniority

Senior, hands-on IC with team leadership responsibilities.

Sourced via workable · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.