Site Reliability Engineer (SRE) (m/w/d)
Core
Design, implement, and operate modern Observability solutions to ensure transparency, stability, and reliability of distributed Cloud, Edge, and Kubernetes platforms.
Role type
Senior Site Reliability Engineer (Observability)
Builds
End-to-End Observability pipelines for Metrics, Logs, and Traces across distributed applications and services.
Domain
Cloud-Native Infrastructure, Distributed Systems, Security Monitoring
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
OpenTelemetry (Collector, OTLP, SDKs, Auto-Instrumentation), Distributed Tracing (Span Correlation, Context Propagation, Sampling), Jaeger (Trace Storage, Indexing, Query Optimization), SIEM Integration (CEF, Syslog), Programming (Go, Python, Shell), Kubernetes Observability, IT Security Awareness
Preferred skills
Metrics/Logs/Tracing pipeline implementation, Correlation mechanisms between Logs/Metrics/Traces, Telemetry/Instrumentation standards for Cloud-Native platforms, Observability in Air-Gap/Edge/Fog environments
Technologies
OpenTelemetry, Jaeger, Kubernetes, Go, Python, Shell, SIEM, CEF, Syslog, OTLP
Responsibilities
Develop and operate pipelines for Metrics, Logs, and Traces; Integrate Security, Infrastructure, and Platform Events into central SIEM solutions; Automate integration and operational processes; Implement and operate Jaeger as a central platform for Distributed Tracing; Develop standardized Telemetry and Instrumentation concepts.
Seniority
Senior, hands-on IC