Senior Site Reliability Engineer - Observability (x/f/m)
Core
Shape observability strategy and ensure platform reliability, debuggability, and scalability for a European-scale healthcare platform serving millions of patients and professionals.
Role type
Senior Site Reliability Engineer (Observability)
Builds
Scalable logging, metrics, tracing, and alerting capabilities for a cloud-native platform.
Domain
Healthcare technology / Cloud Infrastructure / Observability
Deliverable
production ML models | infrastructure
Required skills
Large-scale production platform experience, Cloud platforms (AWS/Azure/GCP), Containerization and orchestration (Docker, Kubernetes), Helm, ArgoCD, Observability tooling (Fluent Bit, OpenTelemetry, Loki, Elasticsearch, Prometheus, Thanos, Datadog), Programming languages (Ruby, Python, Go, Java), Infrastructure as code
Preferred skills
Open-source observability contributions, High-growth tech environment experience, Developer experience passion
Technologies
Kubernetes, Prometheus, OpenTelemetry, Loki, ArgoCD, Ruby, Python, Go, Fluent Bit, Elasticsearch, Thanos, Datadog
Responsibilities
Lead observability strategy across the platform, Identify and lead large-scale cross-cutting reliability initiatives, Take part in the on-call rotation and improve alerting and telemetry
Seniority
Senior, hands-on IC