Site Reliability Engineer (SRE) (m/w/d)
Core
Build and operate a high-reliability monitoring and observability stack for distributed Fog and Edge infrastructures, ensuring stability, performance, and scalability under resource-constrained and disconnected conditions.
Role type
Senior Site Reliability Engineer (Edge/Fog Computing)
Builds
Production-grade monitoring dashboards, alerting strategies, and resilient observability platforms for Fog/Edge nodes
Domain
Cloud Native, Edge Computing, Fog Computing, Observability
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Prometheus (PromQL, Federation, Remote Write, Local Retention), Grafana (Dashboarding, Alerting, Provisioning-as-Code), Loki (LogQL, Retention, Compaction), AlertManager, Kubernetes, YAML, Cloud Native Architecture, Edge/Fog Computing, Resource Optimization (CPU/Memory/Storage), Offline/Air-Gap Operations
Preferred skills
Infrastructure as Code (Terraform, Ansible), GitOps, Linux System Administration, OpenTelemetry, High Availability Infrastructure
Technologies
Prometheus, Grafana, Loki, AlertManager, Kubernetes, Terraform, Ansible, OpenTelemetry, YAML
Responsibilities
Design and implement alerting strategies for autonomous and disconnected operations; optimize monitoring stack performance on resource-constrained Fog nodes; automate monitoring component deployment via Infrastructure-as-Code; develop specialized dashboards for hardware and network monitoring; manage log retention and rotation in edge environments.
Seniority
Senior, hands-on IC