Senior Site Reliability Engineer - Observability (x/f/m)
Core
Shape Doctolib's observability strategy and ensure platform reliability, debuggability, and scalability at a European scale by building logging, metrics, tracing, and alerting capabilities.
Role type
Senior Site Reliability Engineer (Observability)
Builds
Scalable, developer-friendly logging and tracing capabilities; incident detection, response, and postmortem analysis tools.
Domain
Healthcare technology / Cloud-native platform engineering
Deliverable
production ML models | infrastructure
Required skills
Large-scale production platform experience, cloud platforms (AWS/Azure/GCP), containerization and orchestration (Docker, Kubernetes), Helm, ArgoCD, observability tooling (Fluent Bit, OpenTelemetry, Loki, Elasticsearch, Prometheus, Datadog), programming (Ruby/Python/Go/Java), infrastructure as code, troubleshooting complex environments.
Preferred skills
Open-source observability contributions, high-growth tech environment experience, passion for developer experience and platform engineering.
Responsibilities
Lead observability strategy across the platform; identify and lead large-scale cross-cutting reliability initiatives; participate in on-call rotation and improve alerting/telemetry quality.
Seniority
Senior, hands-on IC